Skip to main content

4 min read

Statistics Pitfall #2: Aggregated Data

Statistics Pitfall #2: Aggregated Data

There are plenty of arguments for letting data guide your decisions. “Numbers don't lie” is a phrase you hear often. At first glance, that's entirely true. Data is an unbiased reflection of reality and a good check on assumptions and gut feelings. Unfortunately, it's not always easy to interpret data correctly. It's no coincidence that universities in the Netherlands, almost without exception, fill part of their curriculum with statistics courses. Anyone who works with data runs the risk of falling into one of many pitfalls. In this series of blogs, we'll explain a number of common pitfalls and give you concrete tips to avoid them.

The danger of aggregated data:

On August 4, 2017, the film “An Inconvenient Sequel: Truth to Power” was released, a sequel to Al Gore's hit documentary. About 25 days later, the film had a score of 5.4 out of 2,645 reviews on the well-known Internet Movie Database (IMDB). A middling score that wouldn't necessarily put the film very high on many a movie lover's “must watch” list. What's interesting about this score, however, is that the average rating doesn't seem to tell the whole story. If we look at the distribution of the given scores in a chart, we see the following:

As you can see, opinions about the film were sharply divided, with the scores 1 and 10 given especially often. So if you use the average score to decide whether the film is worth watching, you'll probably get it wrong, since the odds seem almost fifty-fifty that you'll either love it or really hate it. Researchers at ABC News found at the time that there were large differences in opinion between critics and the public, men and women, and even between people who had seen the film and people who hadn't.

The example of Al Gore's film shows the danger of using aggregated metrics. The advantage of such metrics is that it doesn't matter how large your dataset is, you can always summarize it using statistical formulas. But as with any summary, you lose all context in the process. Mathematician Anscombe showed in 1973 that you can construct datasets that, when summarized with aggregated metrics, appear to be exactly the same dataset. He created 4 datasets that all had the same average, but also the same variance and correlation. However, if you visualize them in a chart, the datasets all look very different:

Are aggregations “bad” then? No, not necessarily. Like any other summary, they help you quickly get a general idea and can prompt you to investigate further. In fact, they're sometimes even necessary to form the overall picture. A good example of this is the case of the University of Berkeley. In the 1970s, Berkeley was accused of discriminating based on gender, with women allegedly being structurally admitted less often than men. Researchers looked at the data from the university's 6 largest departments and found the following:

Department Men Women
Applied Admitted Applied Admitted
A 825 62% 108 82%
B 560 63% 25 68%
C 325 37% 593 34%
D 417 33% 375 35%
E 191 28% 393 24%
F 373 6% 341 7%
Total 2691 45% 1835 30%

In the table, the underlined figures per department are the highest in that column for that category. Looking at these 6 departments, it actually seems like the opposite is true. In 4 of the 6 departments, a higher percentage of women than men were admitted. However, if we look at the total, we see confirmation of the claim that men are admitted more often. This effect is also known as Simpson's Paradox. What's happening here is that women applied more often to programs where a lot of people are rejected in general.

As has often turned out to be the case, the way you look at your data is crucial for forming the right insight. The tips below can help you avoid the pitfalls of aggregated data:

#1 Summary

Treat an aggregated metric as exactly that: a summary. Give the user of your dashboard context by visualizing the data.

#2 Totals

Adding a total row to a table makes it easier to recognize an effect like Simpson's Paradox.

#3 Check size

Check your data for sample size. Simpson's Paradox occurs when you have far more observations for one dimension value than for another.

More statistics pitfalls?

Curious about the other blogs in this series after reading this one? Read them all at your leisure via the buttons below.

Pitfall 1Pitfall 3 Pitfall 4

Written by Lennaert van den Brink
Senior Consultant