ML Atlas

11 · Laws · 4 min read · Interactive · updated

What is Simpson's paradox and how do you spot it in data?

In short

A relationship visible in the pooled data can vanish or reverse once the data is split into groups. The 1973 Berkeley admissions case, and why it happens.

What it is

A trend visible in each group separately can disappear or reverse when the groups are combined into one whole. It was described by Edward H. Simpson in 1951 (Karl Pearson and George Udny Yule had noticed it earlier), and its most famous example became the analysis of graduate admissions at Berkeley, published by Bickel, Hammel and O'Connell in Science in 1975.

The paradox is not an arithmetic mistake. Both numbers — the aggregate one and the one broken down by group — are correct. The problem is that they answer different questions, and we usually take whichever one is at hand.

That is why Simpson's paradox is above all a lesson about causality: to know which number to believe, you have to know how the data came about.

Mechanism — why it works this way

An overall rate for a group is a weighted average of the subgroup rates, with the subgroup sizes as weights. If the two groups being compared (e.g. women and men) are distributed across subgroups in very different proportions, you are comparing averages with different weights. It is enough for one group to end up more often in the "hard" subgroups for its overall result to fall, even if it does better in every single subgroup.

In causal language: there is a third variable (here: the department) that affects both who applies and the chance of admission. This is a confounding variable (a confounder). The aggregate comparison mixes the effect of sex with the effect of the choice of department.

Should you always split by groups? No. Judea Pearl shows that the decision depends on the causal structure. If the splitting variable is a cause of both things being compared (a confounder), you need to condition on it. If it is a consequence of the treatment or a mediating variable, splitting on it can introduce bias. The data alone cannot settle this — you need a model of what affects what.

By example

Berkeley, autumn 1973, the six largest departments (A–F). Men submitted 2,691 applications and received 1,198 offers — 44.5%. Women submitted 1,835 applications and received 557 offers — 30.4%. A gap of more than 14 percentage points looks like discrimination.

Now each department separately (women vs men admitted): A — 82.4% vs 62.1%, B — 68.0% vs 63.0%, C — 34.1% vs 36.9%, D — 34.9% vs 33.1%, E — 23.9% vs 27.7%, F — 7.0% vs 5.9%. Women have a higher admission rate in four of the six departments, and in the remaining two the differences are small. The key: 51.5% of men's applications went to the easy departments A and B (which admitted 63–64% of applicants), but only 7.2% of women's. Women applied mainly to departments C–F, which admitted between 6% and 35% of applicants. Had women distributed their applications the way men did, at their own per-department rates they would have had a 51.6% admission rate — higher than men.

In practice

  • Before comparing groups, break the result down by the most important variables: df.groupby(['dept', 'sex'])['admitted'].mean().unstack().
  • Check whether the groups being compared have a similar distribution across subgroups: pd.crosstab(df.sex, df.dept, normalize='index').
  • To adjust, use standardisation (weights from one group) or a model that controls for the variable, e.g. logistic regression admitted ~ sex + dept.
  • Draw a simple causal graph (what affects what) before deciding what to split the data by.
  • The paradox comes back in model evaluation: a model that is better in every segment can have a worse overall score if it was tested on a different mix of segments.

Frequently asked questions

Was there discrimination against women at Berkeley?
These data do not show it at the level of departmental decisions — in four of the six departments women were admitted more often. Bickel and co-authors pointed out that the gap came from where women applied. Why the departments women chose had fewer places is a different question.
Which number should you believe: the aggregate or the per-group one?
It depends on the causal question. If the splitting variable affects both things being compared, believe the per-group numbers. If it is a consequence of the cause under study, splitting on it can distort the result.
How often does Simpson's paradox occur in practice?
A full reversal of direction is rare, but an effect weakening or strengthening after controlling for a confounder is an everyday occurrence in medical, business and social data.

Sources

  • Edward H. Simpson, "The interpretation of interaction in contingency tables", Journal of the Royal Statistical Society, Series B 13(2), 1951, pp. 238–241.
  • Peter J. Bickel, Eugene A. Hammel, J. William O'Connell, "Sex bias in graduate admissions: data from Berkeley", Science 187(4175), 1975, pp. 398–404.
  • Judea Pearl, "Causality: Models, Reasoning, and Inference", 2nd ed., Cambridge University Press, 2009, ch. 6.
  • Judea Pearl, Madelyn Glymour, Nicholas P. Jewell, "Causal Inference in Statistics: A Primer", Wiley, 2016, ch. 1.

See also