04 · Evaluation · 5 min read · Interactive · updated
What is the difference between precision and recall in classification?
In short
Precision is the share of a model's alarms that are real; recall is the share of real cases the model caught. Improving one usually makes the other worse.
What it is
Precision is the share of examples labelled positive by the model that really are positive: precision = TP / (TP + FP). Recall (sensitivity, true positive rate) is the share of actual positives that the model detected: recall = TP / (TP + FN). Both metrics focus on the positive class and ignore true negatives.
Intuition: picture a fisherman with a net. Precision answers the question "how much of what I hauled in is fish rather than rubbish?". Recall asks "what share of the fish in the lake did I catch?". A small net cast only where fish are certain to be gives high precision and low recall. A huge net dragged across the whole lake gives high recall and low precision.
Precision tells you what a single "yes" from the model is worth. Recall tells you how many cases slip through. In medicine recall is called sensitivity, and precision is also known as positive predictive value (PPV).
Mechanism — why it works this way
A classifier usually outputs a continuous score (a probability, a distance from the boundary), and we get a label by comparing it with a threshold. Raising the threshold makes the model say "yes" only in the most confident cases: false alarms decrease (precision rises), but some true positives fall below the threshold (recall drops). Lowering the threshold does the opposite. That is why precision and recall are in a trade-off — every threshold is a different point on this exchange, and the precision–recall curve shows all of them at once.
Recall depends only on how the model treats actual positives, so it does not change when the frequency of the class in the population changes. Precision does. When positives are rare, even a small rate of false alarms among a huge number of negatives exceeds the number of true hits. A test with 99% sensitivity and 99% specificity, for a disease affecting 1 in 1,000 people, has a precision below 10%: for every sick person detected there are about ten healthy people with a positive result. This is the same mechanism as the base rate fallacy.
Why these metrics and not accuracy? With imbalanced classes, accuracy rewards the model for correctly naming the numerous negative class. Precision and recall look only at the class we care about, so an "always no" model gets a recall of zero, even though its accuracy may reach 99%.
Caveat: each metric is incomplete on its own. A recall of 100% is achieved by any model that always says "yes"; a precision close to 100% by a model that says "yes" just once, in its most confident case. Always report them together, along with the threshold at which they were computed.
By example
Logistic regression with standardisation on the Breast Cancer Wisconsin dataset (positive class = malignant tumour; 426 training examples, 143 test examples, stratified split with random_state=0). The test set contains 53 malignant tumours. At the default threshold of 0.5 the model labelled 51 tumours as malignant, and all of them were: precision 100%, recall 96.2% (2 missed cancers).
Lowering the threshold to 0.2: precision 85.2%, recall 98.1% — only one tumour missed now, but 9 benign ones flagged as malignant. At a threshold of 0.1 recall reaches 100%, while precision drops to 82.8% (11 false alarms). In the other direction, at a threshold of 0.7 precision stays at 100% and recall falls to 88.7%: six misses. The model is the same — all that changes is the decision about how much risk of a miss we are willing to trade for false alarms.
In practice
precision_score(y_true, y_pred)andrecall_score(y_true, y_pred);classification_reportprints both metrics for every class.- Decide which label is positive (
pos_label=); in scikit-learn it is 1 by default. - Change the threshold on the probabilities from
predict_proba, or useTunedThresholdClassifierCV(scikit-learn 1.5+), which picks the threshold by cross-validation. - Choose the threshold on validation data, not on the test set — otherwise the reported precision and recall are inflated.
- If one of them matters most, set a requirement for it and optimise the other, e.g. "maximum precision at a recall of at least 95%".
- Typical mistake: reporting precision measured on artificially balanced data (e.g. after SMOTE or 1:1 sampling) as the precision to expect in the population.
Frequently asked questions
- Which matters more: precision or recall?
- It depends on the costs of errors. When a miss is expensive (cancer, fraud, equipment failure), recall is the priority. When a false alarm is expensive (a wrongly blocked account, an important email in spam), precision is the priority.
- How does recall differ from specificity?
- Recall is the share of positives detected; specificity is the share of negatives correctly recognised: TN / (TN + FP). The recall–specificity pair describes the model independently of class frequencies; the precision–recall pair is more sensitive to how the model handles a rare class.
- How do you combine precision and recall into a single number?
- Most often with the F1 score, the harmonic mean of the two. For a threshold-independent assessment, the area under the precision–recall curve (average precision) is used.
Sources
- Géron A. "Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow", 3rd ed., O'Reilly 2022, ch. 3 (Classification — Precision and Recall).
- Murphy K. P. "Probabilistic Machine Learning: An Introduction", MIT Press 2022, ch. 5.1 (Bayesian decision theory — ROC and PR curves).
- Sokolova M., Lapalme G. (2009). A systematic analysis of performance measures for classification tasks. Information Processing & Management, 45(4).
- scikit-learn documentation: Precision, recall and F-measures — https://scikit-learn.org/stable/modules/model_evaluation.html#precision-recall-and-f-measures