ML Atlas

11 · Laws · 4 min read · Interactive · updated

What is the bias–variance tradeoff in machine learning?

In short

A model's error on new data splits into bias, variance and noise. A simpler model has more bias, a richer one has more variance.

What it is

A model's expected squared error on new data equals the sum of squared bias, variance and irreducible noise, and reducing one component usually increases another. The decomposition is a classical result in statistics; it was introduced to neural networks by Stuart Geman, Elie Bienenstock and René Doursat in 1992 under the name "the bias/variance dilemma".

Bias is systematic error: how far from the truth the model is when averaged over many possible training sets. Variance is spread: how much the model changes when it gets a different sample from the same source. Noise is the part of the variation in the labels that no model can predict.

Intuition: a straight line fitted to a sine wave comes out similar every time — and wrong every time (high bias, low variance). A high-degree polynomial hits the wave on average, but each sample produces a different, jittery curve (low bias, high variance).

Mechanism — why it works this way

For a point x and a model f̂ trained on a random training set, the following identity holds: E[(y − f̂(x))²] = (E[f̂(x)] − f(x))² + Var[f̂(x)] + σ². This is not an empirical observation but algebra: add and subtract the average prediction, and the cross terms vanish in expectation.

The trade-off arises because both components depend on the model's flexibility in opposite directions. A rigid model cannot express the true relationship, so its average is wrong. A flexible model can express almost anything — including the random noise of a particular sample — so what it ends up expressing depends on the luck of the sample.

Variance is lowered by more data (noise averages out), regularisation (penalising extreme parameters) and averaging many models (bagging, random forests). Bias is lowered by a richer model class or better features. Nothing reduces noise except better measurements.

An important caveat: the picture of a "U-shaped curve" as a function of the number of parameters is not a law of nature. The error decomposition is always true, but variance need not grow with the number of parameters — in heavily overparameterised models it can fall again. This is the double descent phenomenon, which forced a more careful definition of what "complexity" actually measures.

By example

We took 25 evenly spaced points x in the interval [0, 1], with labels y = sin(2πx) plus noise with a standard deviation of 0.3 (so the noise contributes σ² = 0.09). A thousand times we drew fresh noise and fitted polynomials of various degrees, then computed squared bias and variance on a grid of points.

Degree 1 (a straight line): squared bias 0.186, variance 0.007, total error 0.28. Degree 3: 0.005 and 0.012, total 0.107. Degree 5: bias practically zero, variance 0.017, total 0.108. Degree 9: variance 0.031, total 0.121. Degree 13: variance 0.052, total 0.142. Bias is gone by degree 5, and from then on each additional degree only adds variance. The best compromise is degree 3–5.

In practice

  • Diagnosis is done with curves: validation_curve (error vs complexity) and learning_curve (error vs number of examples).
  • A large gap between training and validation scores → high variance: more data, regularisation (Ridge, alpha, max_depth, dropout), bagging.
  • Both scores poor and close together → high bias: a richer model, better features (PolynomialFeatures), less regularisation.
  • Random forests (RandomForestClassifier) are a practical machine for reducing the variance of trees without raising their bias.
  • Typical mistake: choosing complexity on the test set — the measured variance is then understated.

Frequently asked questions

Can you have low bias and low variance at the same time?
Yes, if there is a lot of data relative to the complexity of the problem, or if the model makes good assumptions about the structure of the data. The trade-off is strict only for a fixed amount of data and a fixed family of models.
Does this trade-off still hold in deep learning?
The error decomposition always holds, because it is a mathematical identity. What does not hold is the naive rule "more parameters = more variance" — very large networks often generalise better than medium-sized ones.
How does model bias differ from bias in the sense of prejudice (fairness)?
They are different concepts that share a name. Here bias is the systematic error of the prediction relative to the true function; in the fairness context it means unequal treatment of groups.

Sources

  • Geman S., Bienenstock E., Doursat R. (1992). Neural Networks and the Bias/Variance Dilemma. Neural Computation, 4(1), 1–58.
  • James G., Witten D., Hastie T., Tibshirani R. (2021). An Introduction to Statistical Learning, 2nd ed. Springer, ch. 2.2.2.
  • Hastie T., Tibshirani R., Friedman J. (2009). The Elements of Statistical Learning, 2nd ed. Springer, ch. 7.3.
  • Belkin M., Hsu D., Ma S., Mandal S. (2019). Reconciling modern machine-learning practice and the classical bias–variance trade-off. PNAS, 116(32), 15849–15854.

See also