04 · Evaluation · 4 min read · Interactive · updated
What is a confusion matrix and how do you read it?
In short
A confusion matrix counts how many examples of each true class the model assigned to each class. Accuracy, precision, recall and other metrics follow from it.
What it is
A confusion matrix is a table whose rows correspond to the true classes, whose columns correspond to the classes predicted by the model, and whose cells hold the number of examples with each combination. The diagonal holds the hits; everything off it is a specific kind of error. It is the most basic and most information-rich way to describe a classifier's results.
In binary classification it has four cells: TP (true positives — a sick patient recognised as sick), TN (true negatives — a healthy one as healthy), FP (false positives, a type I error — a false alarm) and FN (false negatives, a type II error — a miss). A single number such as accuracy merges these four values into one and loses the information about what kind of errors the model makes.
Mechanism — why it works this way
Every popular classification metric is a function of these four numbers. Accuracy = (TP + TN) / all. Precision = TP / (TP + FP) — what share of alarms are real. Recall (sensitivity) = TP / (TP + FN) — what share of real cases we detected. Specificity = TN / (TN + FP). Reading the matrix, you see immediately which of these is the problem, instead of guessing from a single number.
The matrix also shows why accuracy can mislead. When 1% of cases are positive, a model that always says "negative" has 99% accuracy and TP = 0. The matrix reveals this at once: the whole "positive" column is empty.
Importantly, a confusion matrix is not a property of the model alone, but of the model and the decision threshold. A classifier usually outputs a probability, and we get a label by comparing it with a threshold (0.5 by default). Lowering the threshold moves examples from the "negative" column to the "positive" one: FN falls, FP rises. No threshold improves everything at once — it is a trade of one kind of error for another, and the exchange rate is set by the costs of errors in the application at hand.
The matrix also depends on the class proportions in the set on which it is computed. The same model capabilities will give a completely different precision in an oncology hospital than in a screening programme, because the share of sick patients changes. Recall and specificity do not depend on this; precision and accuracy do.
In multiclass problems the matrix is K × K. The off-diagonal cells show which classes get confused with each other, which often hints at what the model is missing (e.g. the digits 3 and 8 differ in their left half).
By example
On the Breast Cancer Wisconsin dataset (569 tumours, positive class = malignant) I trained logistic regression with standardised features on 426 examples and evaluated it on 143 test examples (stratified split, test_size=0.25, random_state=0). The test set contained 53 malignant and 90 benign tumours. At a threshold of 0.5 the matrix looks like this: TN = 90, FP = 0, FN = 2, TP = 51. Accuracy is 98.6%, precision 100%, recall 96.2%: the model raised not a single false alarm, but missed two cancers.
In diagnostics a miss is usually more dangerous than a false alarm, so let us lower the threshold to 0.1. The matrix: TN = 79, FP = 11, FN = 0, TP = 53. All malignant tumours are detected, but 11 benign ones would be sent for further tests; accuracy has dropped to 92.3%. Which matrix is better depends on the costs — and that is exactly why it pays to look at the matrix rather than just at accuracy, which would point to the first option.
In practice
confusion_matrix(y_true, y_pred)fromsklearn.metrics; in scikit-learn rows are true classes and columns predicted ones. For the binary casetn, fp, fn, tp = confusion_matrix(...).ravel().- Visualisation:
ConfusionMatrixDisplay.from_estimator(model, X_test, y_test)orfrom_predictions. normalize='true'divides by row sums (the diagonal then shows each class's recall),normalize='pred'by column sums (precision).- Check which class is "positive". In
load_breast_cancerlabel 1 means a benign tumour, so it has to be flipped if cancer is to be the positive class. - Typical mistake: textbooks and libraries use different axis conventions. Always label the axes.
Frequently asked questions
- Why is it called a "confusion matrix"?
- Because it shows what the model confuses each class with. The cell in row "cat" and column "dog" tells you how many cats were taken for dogs, and the opposite cell how many dogs were taken for cats. These two numbers are usually not equal, so the matrix is not symmetric.
- Which error is worse: FP or FN?
- That depends on the application, not on statistics. In screening it is usually FN; in a spam filter rather FP (a lost important email). It is worth settling the costs before training and choosing the threshold to match.
- How do you read a matrix with many classes?
- First look at the row-normalised diagonal — that is each class's recall. Then look for the largest off-diagonal cells: they point to the pairs of classes the model confuses most often.
Sources
- Géron A. "Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow", 3rd ed., O'Reilly 2022, ch. 3 (Classification).
- James G., Witten D., Hastie T., Tibshirani R. "An Introduction to Statistical Learning", 2nd ed., Springer 2021, ch. 4 (Classification).
- Fawcett T. (2006). An introduction to ROC analysis. Pattern Recognition Letters, 27(8).
- scikit-learn documentation: Confusion matrix — https://scikit-learn.org/stable/modules/model_evaluation.html#confusion-matrix