ML Atlas

11 · Laws · 4 min read · Interactive · updated

What does Goodhart's law say and why does optimising a metric break a model?

In short

When a measure becomes a target, it ceases to be a good measure. In ML, optimising a proxy metric pulls it apart from what we actually care about.

What it is

When a measure becomes a target, it ceases to be a good measure. This concise form of the principle was coined by the anthropologist Marilyn Strathern in 1997, summarising a 1975 observation by the economist Charles Goodhart: any observed statistical regularity tends to collapse once pressure is placed on it for control purposes.

In machine learning we almost always optimise a proxy: log loss instead of the business decision, a benchmark score instead of general ability, a reward model's rating instead of a person's true preference. As long as optimisation is weak, the measure and the goal move together. The harder we optimise, the more likely we are to find a way of improving the measure that does not improve the goal — or even harms it.

Classic ML examples: a reinforcement learning agent that, instead of winning a race, goes round in circles collecting bonus points; a language model fine-tuned against a reward model that learns long, ingratiating answers because they get higher scores.

Mechanism — why it works this way

A measure is the goal plus error: proxy = goal + ε. When we pick the solution with the highest proxy out of many candidates, we are selecting both high values of the goal and high values of the error. The more candidates we search (stronger optimisation), the larger the share of the measured gain that comes from the error rather than the goal. This is the effect of selection on a noisy measurement, the same as the winner's curse.

David Manheim and Scott Garrabrant (2018) distinguished several variants. The regressional variant: selection on noise alone means the chosen candidates are worse than the measure suggests. The extremal variant: at extreme values, the relationship between the measure and the goal can look quite different from the typical range. The causal variant: we optimise something that correlated with the goal but did not cause it. The adversarial variant: someone (or the model itself) deliberately exploits a loophole in the measure.

When the measurement error has heavy tails — rare but enormous mistakes — strong optimisation selects almost nothing but those mistakes. Gao, Schulman and Hilton (2023) measured this for reward models in language model training: as optimisation strength increases, the reward model's score keeps rising, while true quality first rises and then falls.

By example

A simulation: N candidates, each with a true quality drawn from a normal distribution, measured with error. We pick the candidate with the highest measured score. When the error is normal (on the same scale as quality), with 10 candidates the chosen one has a measured score of 2.17 and a true quality of 1.09; with 100 — 3.55 and 1.77; with 10,000 — 5.45 and 2.71. The true gain grows, but only by half of what the measure promises.

When the error has heavy tails (Student's t distribution with 2 degrees of freedom), the picture flips. With 10 candidates: measure 4.0, quality 0.72. With 100: measure 12.4, quality 0.35. With 1,000: measure 40, quality 0.17. With 10,000: measure 119, quality 0.05 — almost zero. The harder we optimise, the higher the measure and the worse the actual choice. This is Goodhart's law in its purest form.

In practice

  • Track several metrics at once (e.g. roc_auc, average_precision, calibration) and check whether an improvement in one comes at the cost of the others.
  • Keep a held-out test set on which you optimise nothing, and check it rarely.
  • In RLHF, a KL penalty relative to the base model is added to limit how hard the reward model is optimised.
  • Regularly look at concrete examples of the model's outputs — a number will not reveal that the model is "gaming" the metric.
  • When a metric starts rising suspiciously fast, look for a loophole in the metric first, and celebrate only afterwards.

Frequently asked questions

Does Goodhart's law mean we should not use metrics?
No. It means a metric is a diagnostic tool, not an end in itself. The harder you push on it, the more often you need to check that it still measures what it should.
How does it relate to leaderboard overfitting?
It is the same mechanism. The score on a public leaderboard is a proxy for quality on new data; repeatedly optimising for it fits the model to the leaderboard's noise.
How does Goodhart's law differ from Campbell's law?
In 1979 Donald Campbell formulated a similar principle for social indicators: the more an indicator is used for decision-making, the more it is subject to pressures that distort the very process it was meant to measure. Both principles describe the same problem from different fields.

Sources

  • Goodhart C. A. E. (1975). Problems of Monetary Management: The U.K. Experience. Papers in Monetary Economics, Reserve Bank of Australia.
  • Strathern M. (1997). "Improving Ratings": Audit in the British University System. European Review, 5(3), 305–321.
  • Manheim D., Garrabrant S. (2018). Categorizing Variants of Goodhart's Law. arXiv:1803.04585.
  • Gao L., Schulman J., Hilton J. (2023). Scaling Laws for Reward Model Overoptimization. ICML 2023.
  • Campbell D. T. (1979). Assessing the Impact of Planned Social Change. Evaluation and Program Planning, 2(1), 67–90.

See also