Thursday, February 19, 2009

The Practical Statistician-A Toolkit

I have had the pleasure of working with a lot of statisticians, mathematicians, data miners and econometricians (let's call them PEMD-persons extracting meaning from data, for ease) in my career. An observation I have often made is that while all of them know the tools of their trade, only a few eventually go on to become excellent practitioners or as I call them 'practical statisticians' in the industry. What is it that these experts have that gets them far ahead in their trade? A toolkit that helps them survive the real world journey. Here is the list of items in that toolkit:


Item #1: Pen and notebook (a thick one)-they carry this around at all times even to bed. This helps them make copious notes when others are talking and think aloud when they are structuring their thoughts, attacking problems and analyzing outputs. They guard this notebook zealously and get visibly upset if it ever gets lost or misplaced. They recognize that in order to streamline loads of work, manage their time well, analyze the problem fully and present the output lucidly without going insane they must structure their thoughts. Written matter is the key.

Item #2: Three books for reference and speed reading skills-one is usually about the software they are using, the other two are the best applied texts on most used techniques in their field and new emerging areas(which no one else has a clue about). They read many more research articles than other people (and yes they usually do that during their breaks or in their leisure time). If they don’t understand an article the first time round, they absolutely have to read it again and again till they do.

Item #3: Data dirty fingers-they execute projects no matter how high they rank in the corporate hierarchy. They recognize that leading from the front means ability to do the work at the back end especially when all hell breaks lose.

Item #4: Non-technical speak-they are able to communicate their ideas and statistical methods to a wide audience without using statistical jargon.

Item #5: Graphs-they like to graph data and get a sense of numbers visually. This ability to look at both numbers and graphs helps them get a finer sense of the data and what they don’t know and must find out from it.

Item #6: A good dose of imagination, critical thinking and skepticism-they function like detectives and for them most business problems present cases to be cracked. After the project starts they devote all effort in cracking the case oblivious to everything and everyone else.

Item #7: Mentoring and training calendar-unless they pass on their wisdom and how they put the problem, method and experience together, they know they will continue to do the same work over and over again.

Item #8: A broad view of their role -they define their role rather than let client's, coworkers and organizations peg them. They like their roles to be larger and ‘more whole’ not constrained by their degree and specialization.

Item #9: Practical adequate solutions-while striving for the best solution, they recognize that they may need to deliver less optimum solutions based on project constraints and client readiness.


Item #10: A passion for statistics-especially it's applications in different fields, and an understanding of what it can and cannot do.

Wednesday, February 18, 2009

Trends: R vs. SAS-What's really at the heart of the matter?

Okay, I promised myself that I would not jump into this debate and I bit my tongue and fingers like a thousand times last week. Go ahead and shoot me I'm only human.

Here goes...

Methinks this R vs. SAS debate is less about the merits and demerits of the two software and more about the David vs. Goliath(or Hare vs. Tortoise) effect. David, in this case also provides strong competition in a slightly monopolistic market situation.



I have worked with SAS and I don't have a strong opinion against it(except for it's really bad graphics). I am new to R and I like it(yes, there will be some pet peeves as time goes by). I have also used most other competing software in this space(SPSS, Minitab, Stata, Matlab etc).


So what's the issue you may ask? Well, no matter what anyone says(or posts), I believe one of the main reasons that R is generating such a lot of press(and don't get me wrong-it has strong merits) is the fact that with all it's merits it is also FREE! Whether we like to admit it or not, it bothers us that we have to pay for using SAS when R which is as good(if not better in some areas) is available for zero cost. Would the same debate be as heated if R did not deliver? I doubt it.

Add to this the point that R comes in as the 'underdog' that most of us like to see win and you get a better idea of why there is so much angst all around on this issue.

Enough said.

Tuesday, February 17, 2009

Parlez vous Statistics?

As I sat debating the issue about whether we should have an informal case study 'test' for statisticians who want to work or intern with us, I read Andrew Gelman's blog article on a new course in statistical communication that he would like to teach sometime. It brought home to me the fact that if I went ahead with this test we would not have any new hires at all, since most would flunk out.

Why oh why do we not teach statistical communication at most universities or even at jobs? The lack of this skill has made proponents and users of the subject in the industry unable to communicate in the same language.

So what are some of the skills I would test for statisticians? Here is my list :
  1. Translating a business problem to an analytical and statistical problem
  2. Writing a proposal(or at least the proposed analytical solution part of the proposal)
  3. Creating a process/flow chart of the analytical solution
  4. Graphical presentation of data(raw, cleaned and analysed or modeled)
  5. Summarizing and communicating statistical results in both technical and non-technical ways(depending upon the audience). This would also include documentation of the project and an executive summary of findings
  6. Ability to write simple and elegant computer code and read the same(irrespective of software and writer differences)
  7. Collaborative work effort with other colleagues(programmers, consultants, academicians etc)
  8. Knowledge of statistical pitfalls
  9. Other communication skills(e-mails, blogs, discussions, knowledge sharing etc)
  10. Ability to read, understand and summarise research papers(good knowledge of work in relevant focus areas)
More organisations need to get involved with universities to encourage teaching of these skills at an early level to practitioners who plan to join the industry. Statisticians on the other hand need to move out of their comfort zone and ensure that they become adept at communicating their language to a wider audience. Once the above skills have been mastered along with a sound knowledge of statistics, we may finally be viewed as the geeks with the sexy job(as Hal Varian-Google's chief economist points out in an interview with the McKinsey Quarterly).

Wednesday, January 28, 2009

Plan a coming out party for Outliers-Part 1

I picked up Malcolm Gladwell's book on Outliers because it got me thinking about an issue I encounter in all my analyses. I remembered the case quoted by statisticians as reason for not leaving out outliers. For those who don't already know the story, here it is:

In 1985 three researchers-Farman, Gardinar and Shanklin were extremely puzzled by data gathered by the British Antarctic Survey showing that ozone levels for Antarctica had dropped 10% below normal January levels. The reason for the puzzlement was because the Nimbus 7 satellite, which had sophisticated instruments aboard for recording ozone levels, hadn't recorded similarly low ozone concentrations.


When they examined the data from the satellite they realised that the satellite in fact had been recording these low concentrations levels and had been doing so for many years. Because the ozone concentrations recorded by the satellite were so low, they were being treated as outliers by a computer program and left out of the analysis.


The Nimbus 7 satellite had in fact been gathering evidence of low ozone levels since 1976. Due to the outliers being discarded without being examined, the damage to the atmosphere caused by CFC's went undetected and untreated for up to nine years(this account is disputed by NASA researchers who say that they had flags in place for low values and did notice the low ozone values and subsequently presented thier paper but the Farman trio's paper on the same beat them to it).

What is the moral of the story? To take a deeper look at outliers in your data because they usually tell a unique story if you are really willing to listen.

Why should you be looking for outliers-you may ask, here's why:

  1. Erroneous results in reporting, dashboards and executive summaries as these are comprised mainly of 'mean/average' numbers. Statistical tests and analysis may be negatively affected.
  2. The outlier may be the story of interest in your data i.e. the high value accounts, the seasonal spenders, the defaulters etc.
  3. The understanding that an outlier may yield about the data gathering process. I remember years ago being asked to cross check data when the results showed that median age of women at the birth of their last child was mid forties for certain eastern Indian states. This result was completely off from the national average(which was lower) and the client suspected a data issue at the agencies end. On enquiry, we learnt that women in these states were losing teenage children(that they had when they were much younger) to terrorism and drugs and thus were having more children in later years. In this case the explanation held else we would have had to investigate why the error occurred in the data collection.

Once you find outliers in the data, what do you do ? Before anything else-report them! It does not matter if they are few in number, if you understand why they occurred or if you plan to leave them out for whatever good reason. In a lot of analysis, I see a disturbing trend of suppressing or 'fixing' outliers without understanding them or reporting them.

My suggestion thus is to have a discussion on outliers in your data before deciding what you will do with them. I will talk about addressing outliers in part 2 of this piece, but for now here are some things to mull over-

  • Can tracking financial performance of companies and individuals and identifying outliers help curb scams of the Satyam and Maddoff type? Markopolos and mathematician DiBartolomeo warned regulators for years that Madoff could not be consistently generating higher than market profits unless he was running a ponzi scheme. I am sure some people out there also looked through Satyam's records and had misgivings but kept quiet.

  • Last year Republican representative Mark Souder proposed that baseball players whose on-field statistics suddenly improved should be tested more often for performance-enhancing substances. The thought is to measure actual player performance against projected performance and history based on a typical career path and identify outlier performance or sharp deviance. Maybe undertaking this analysis for track athletes may give sharper results.

What does all this have to do with Gladwell's book on outliers? Nothing really, except a reiteration to take a fresh and deep look at outliers before tossing them away or standardising them. As for the book, it was alright(not earth shattering), the reason behind the Korean airline crashes made for the best reading.

Wednesday, January 7, 2009

Book Review: Super Crunchers-The fight between experts, gut and data

I enjoyed reading Ian Ayres book. Let me say that again and right - I really enjoyed reading Ian Ayres book. For those who have not already read the book-it details how data driven number crunching algorithms work better than expert predictions and gut feelings and how super crunchers(read-statistically literate and number crunching savvy) individuals will have an edge in decision making in the future.

I liked the book because it lucidly illustrates trends that I have seen in the last decade-a better adoption of predictive models among businesses, more data generation and storage, an industry wide need for talented number crunchers and the conflict when data driven approaches come face to face with the resident expert or the manager who swears by his gut.

The case studies are very interesting and apt-it was amusing to read about the prediction of a vintage by an algorithm(I must pick up some wine based on the prediction soon). I could empathise with the story about a fellow economists frustration at waiting to get the final odds number on the Downs syndrome screening for his unborn child, and the inability of the technicians to apply the Bayes theorem(I've been there). As Ayres points out, I agree neural nets have a long way to go before they replace other mainstream techniques and it's not just due to the over fitting problem. Randomised trials still need to become mainstream among most marketers.


What really makes the book stand out is that data crunchers like me along with millions others 'get it'. I build predictive models that are elegant and simple and able to help clients make better decisions about their businesses. We constantly face sceptics about how predictive models can fare better than the resident experts knowledge of his market or brand or business. We sometimes pitch to client's who tell us their business problems cannot be put in an equation(it makes me squirm because I have a personal data project on which aims at predicting market prices for Indian contemporary art). After years in statistics, it's still difficult to help people understand standard deviation or 2SD.

Do I agree with the book's central premise-yes I do. In a data driven world, let numbers do the talking-stand aside experts and intuition.

Friday, December 5, 2008

Trends: Going the way of R and other open source software

My colleague Girish recently mailed me a New York Times business computing article about how data analysts have taken to R as the open source programming language.

The article took me back in time to 1996, when I was a graduate student in the US. Fellow statisticians were raving about R as the new generation data crunching language and something that was going to give other data packages a run for their money. We were at that point doing our statistical number crunching on student licenses of SAS. While intrigued with the whole issue, I was too busy learning applied statistics and SAS and just getting through grad school semesters.

Now years later it is with a feeling of deja vu that I read the article because today I am much closer to embracing R as 'the' crunching language for myself and our business.

We've done the testing and it's won hands down every time;



  • Ease of use


  • Readily available code modules(learning from others is a key here-we techies love to outsmart each other)


  • Wonderful graphics


  • Excellent data manipulation


  • No fees


  • Ability to customise


  • Lots more...


While competitors are quick to dismiss it, R works because it has created a democratic community of statisticians and others who like to see number crunching become easier and more visual. The fact that it is open source provides the added kick to be able to create customised modules that the community can use. It blends programming and statistical skills together more elegantly than I have ever seen. The fact that it has a fan following among my tribe is therefore not surprising.

Thus, is R and other open source software the way to go-absolutely! The reasons are many but let me rank order them based on how we took the leap-

  1. Stacks up and beats competition on most data crunching modules.
  2. Easy to use.
  3. Collaborative value model: the conviction that a collective community can create better thought and tools than a competitive one.
  4. Better service: less downtime, quicker error resolution and a help desk of people dedicated to fixing issues.
  5. Excellent customisation options: The ability to create what you want for your business and put it out there.
  6. Cutting edge graphics.
  7. The geek factor-the thrill of creating, bettering and showing off to other like minded individuals cannot be underestimated.
  8. Lower technology cost: while this is great, believe me this is not the main reason that businesses use open source.




Monday, December 1, 2008

Segmentation-making it more science than art

The reason for delay in posting this has been because I've been toying with whether I should write on segmentation or not. So much has been written on this subject that it makes me a little hesitant about revisiting this space.


What got me to finally pen this was the title of a paper at an upcoming conference that said 'How statistics get in the way of actionable segmentation'. I don't know what the presenters have to say (must source a copy after the conference) but the title made me laugh. The two words that stuck in my head were 'statistics' and 'actionable segmentation' and whether the twain will ever meet.


I have undertaken enumerable segmentation projects in my career, some simple, others complex, yet others that go nowhere. All of them in the end have the same things in common:

  1. Too many bases variables

  2. Over reliance on cluster analysis as the primary tool for segmentation

  3. Use of subjective judgement to evaluate results of the cluster solution

  4. Lack of reliability and validity tests on the solution

  5. Recreation of the scientific solution into a more 'creative and arty' one

What the above means for managers who implement segmentation solutions is that they could be formulating strategy and targeting segments that are unstable and unreal. There exists a body of research that calls for a deeper look at the statistics and data that go into cluster analysis and segmentation(I will be happy to provide the references).

The real issue continues to be an inability of both analysts and practitioners to put together a common road map for segmentation that takes into account statistical robustness of the technique along with creation of actionable segments that can be targeted through focused marketing programs. In my experience, the science of segmentation gets lost in the art.

Here are the five key things that analysts and practitioners must do to create better, robust and more scientific segments-

  1. Choose bases variables for segmentation that tie in with the end goal of segmentation and keep their number not more than 8-10. Build a set of good profiling variables that tie into the bases variables(there is no restriction in number here).
  2. Explore other tools for segmentation(sometimes simple business rules work just as well). Latent class analysis offers excellent alternatives for both survey and crm data and is still a highly underused technique. Try two techniques, if possible and compare and contrast results.
  3. Use a variety of statistical parameters to evaluate a solution instead of relying on one or two or on subjectivity. Decide which metrics you want to look at before the study. For example-dendograms, change rate, psuedo Rsq, hotelling's Tsq can be some metrics for evaluating no. of clusters in a cluster solution. The BIC, p-value, parsimony(no. of parameters) and the bootstrap p-value can be the parameters to nail number of segments in a latent class segmentation. Reliance on 'many' statistics vs. 'few' should be the mantra.
  4. Test reliability through hold out samples and validity through looking at profiling variables and how they differentiate the solution. The holdout sample results must match those of the developmental sample in terms of the number of segments and profiles. Most of the picked profiling variables must adequately differentiate the final segment solution. If there is an issue with reliability and validity-the solution may have a problem. Going back and reworking the same is the best way out.
  5. Don't use the art of segmentation to sidestep the science for a solution, use it if you will to add to the same.