11 · Laws · 4 min read · Interactive · updated
What is Anscombe's quartet and why do you need to plot your data?
In short
Four datasets with identical means, variances, correlation and regression line look completely different when plotted. Why summary statistics are not enough.
What it is
The same descriptive statistics can describe completely different data, so before you trust the numbers, you have to look at a plot. Francis Anscombe showed this in 1973 by constructing four small datasets of 11 points each that have nearly identical means, variances, correlation and regression line.
Plotted, the first dataset is an ordinary cloud of points around a line. The second is a smooth parabola. The third is a perfect line with a single outlier. The fourth is a vertical column of points plus one lone point far to the right. Four different stories, one table of numbers.
Anscombe wrote at a time when computers calculated statistics but drawing plots was expensive, and many statisticians considered graphs less "rigorous" than numbers. The quartet was an argument against that belief, and to this day it is the first example taught in exploratory data analysis.
Mechanism — why it works this way
The mean, variance and correlation coefficient are summaries: each squeezes many numbers into one. Compression by definition loses information, and many different datasets share the same summaries. Pearson's correlation coefficient measures only the strength of a linear relationship and is sensitive to single extreme points, because it is built on squared deviations.
The least-squares regression line depends only on the means, variances and covariance. If these are equal, the line is the same — regardless of whether the data is linear, curved, or dominated by a single point. The method will always fit something; it will not tell you the model is wrong.
That is why the same number (e.g. r = 0.82) can mean "a moderately strong linear relationship with noise", "a perfect non-linear relationship described by the wrong model" or "no relationship at all, just one leverage point". Only a plot of the data and a plot of the residuals settle it.
In 2017 Justin Matejka and George Fitzmaurice generalised the idea: using simulated annealing they generated a dozen datasets (including one shaped like a dinosaur) with statistics identical to two decimal places. So the quartet is not a curiosity but the rule: there are few statistics and infinitely many possible shapes.
By example
Computed from the quartet file: in each of the four datasets the mean of x is 9.0, the variance of x 11.0, the mean of y 7.50, the variance of y between 4.12 and 4.13, the correlation 0.816–0.817, and the regression line is y ≈ 3.00 + 0.500·x with R² ≈ 0.67. From the table alone they cannot be told apart.
The differences emerge on closer inspection. In dataset II a parabola (a 2nd-degree polynomial) explains practically 100% of the variance — the straight line was the wrong model. In dataset III ten points lie exactly on a line with slope 0.345 (correlation 1.000), and a single point with y = 12.74 raises the slope to 0.50 and lowers the correlation to 0.82. In dataset IV ten points share the same x = 8, and the only point with x = 19 single-handedly determines the whole slope — without it, the relationship cannot even be computed.
In practice
- Always plot:
sns.scatterplotfor pairs,sns.pairplot(df)for many variables,sns.lmplot(data=df, x='x', y='y', col='dataset')for the quartet. - After fitting a regression, inspect the residuals (
sns.residplot); a pattern in the residuals means the model has the wrong shape. - Check the influence of individual points: leverage and Cook's distance (
statsmodels→get_influence().cooks_distance). - Alongside Pearson's correlation, compute Spearman's (
df.corr(method='spearman')), which is less sensitive to outliers. - In pandas,
df.describe()is a good start, but never the end of an analysis.
Frequently asked questions
- What does Anscombe's quartet show?
- That means, variances, correlation and the regression line can be identical for data of completely different shapes. Descriptive statistics are no substitute for a plot.
- Does a correlation of 0.82 mean a strong linear relationship?
- Not always. In the quartet the same correlation describes noise around a line, a parabola, a line with one outlier, and data in which the relationship is created by a single point.
- Where can I find the quartet data?
- It ships with many libraries, including seaborn (`sns.load_dataset('anscombe')`) and R (`datasets::anscombe`). It consists of four datasets of 11 (x, y) pairs each.
Sources
- Francis J. Anscombe, "Graphs in Statistical Analysis", The American Statistician 27(1), 1973, pp. 17–21.
- Justin Matejka, George Fitzmaurice, "Same Stats, Different Graphs: Generating Datasets with Varied Appearance and Identical Statistics through Simulated Annealing", Proceedings of CHI 2017, ACM.
- John W. Tukey, "Exploratory Data Analysis", Addison-Wesley, 1977.
- Gareth James et al., "An Introduction to Statistical Learning", 2nd ed., Springer, 2021, ch. 3 (Linear Regression).