03 · Supervised · 5 min read · Interactive · updated
How does a random forest work and why is it better than a single tree?
In short
A random forest averages hundreds of decision trees, each trained on a random data sample and random feature subsets. That makes it accurate and stable.
What it is
A random forest is an ensemble of many decision trees, each trained on a different random sample of the data and choosing, at every split, from a random subset of features. The forest's prediction is the majority vote of the trees (classification) or the average of their predictions (regression). The method in its current form was described by Leo Breiman in 2001.
A single deep tree is like an expert with a perfect memory and poor judgement: it remembers every training case, but its rules change dramatically with a small change in the data. A forest is a council of hundreds of such experts, each of whom saw slightly different data and looked at slightly different features. Their individual quirks cancel out and the shared signal remains.
For two decades the random forest has been one of the safest choices for tabular data: it works well without tuning, does not need feature scaling and rarely fails spectacularly.
Mechanism — why it works this way
The forest combines two sources of randomness. The first is bagging: each tree gets a bootstrap sample — as many examples as the dataset has, but drawn with replacement, so some repeat and about 37% do not make it into the sample at all. The second is feature sampling: at each split the tree considers only m randomly chosen features instead of all of them (for classification typically m ≈ √p).
Why is the second source needed? Averaging reduces variance only when the trees make different errors. The variance of the average of B trees, each with variance σ² and pairwise correlation ρ, is ρσ² + (1 − ρ)σ²/B. The second term vanishes for large B, but the first remains — the correlation between trees sets a floor the forest cannot go below. If one feature is very strong, all trees built by bagging alone will start with it and end up similar. Feature sampling forces some trees to find other routes, which lowers ρ.
Trees in a forest are usually grown without pruning. Each one alone overfits (low bias, high variance), but averaging removes the variance and leaves the low bias. That is why adding trees does not cause overfitting — the score stabilises and the only cost is computation time.
A bonus from bootstrapping: each example was left out of about 37% of the trees, so it can be evaluated using only those trees. This out-of-bag (OOB) error is a free approximation of cross-validation.
Limitations: a forest does not extrapolate (in regression it cannot predict values outside the training range), loses the readability of a single tree and can be large in memory. The built-in feature importance (mean decrease in impurity) favours continuous and high-cardinality features; permutation importance is more reliable.
By example
Breast Cancer Wisconsin: 569 tumours, 30 features; training on 426, testing on 143 (stratified split, random_state=0). A single fully grown tree is correct on 90.2% of test cases, and depending on the random seed (20 different random_state values) anywhere from 89.5% to 93.0%. A forest of 500 trees: 94.4% on test, while the OOB estimate says 96.5% — on its own, without holding out any data. A fairer comparison comes from repeated cross-validation (5 folds × 10 repeats): tree 92.5%, bagging with 100 trees 95.5%, random forest 95.9%.
Feature sampling really does "decorrelate" the trees: the average correlation of predicted probabilities between pairs of trees is 0.83 with bagging alone (all 30 features at every split) and 0.79 when sampling √30 ≈ 5 features. One tree of the "forest" gives 88.1%, five trees already 95.8%, and from 50 trees onward the score sits at 94.4% — the small fluctuations with only a few trees are the noise of 143 test cases. According to the built-in importance, the most important features are worst perimeter (0.15), worst number of concave points of the contour (0.12) and worst radius (0.12) — three features that largely measure the size and irregularity of the tumour.
In practice
RandomForestClassifier(n_estimators=500, n_jobs=-1, random_state=0)andRandomForestRegressor; no feature scaling needed.n_estimators: more means more stable (typically 300–1000); it does not cause overfitting, it only costs time.max_featuresis the main parameter:"sqrt"by default for classification and1.0(all features) for regression; values of 0.3–0.5 are worth trying.oob_score=Truegives a free accuracy estimate; for feature importance usepermutation_importanceon validation data.min_samples_leafof 1–5 smooths predictions and shrinks the model;max_depthis usually left unlimited.
Frequently asked questions
- Can a random forest overfit?
- Adding trees does not make a forest overfit — the error stabilises at a level set by the correlation between trees. A forest can, however, fit noise through overly deep trees on very noisy data; a larger `min_samples_leaf` helps there. 100% training accuracy is normal for a forest and does not indicate a problem.
- Random forest or gradient boosting?
- Gradient boosting (XGBoost, LightGBM) usually achieves slightly better results once tuned, but is more sensitive to hyperparameters. The forest is more robust and good out of the box. In practice the forest is an excellent first model, and boosting is the tool for squeezing out the last few percentage points.
- How do you interpret a random forest?
- You can no longer read a single tree, but you can probe the model from the outside: permutation importance shows which features are needed, partial dependence plots show how the prediction changes with a feature's value, and SHAP values decompose an individual prediction into feature contributions.
Sources
- Breiman L. "Random Forests", Machine Learning 45(1), 2001.
- Hastie T., Tibshirani R., Friedman J. "The Elements of Statistical Learning", 2nd ed., 2009, ch. 15.
- James G., Witten D., Hastie T., Tibshirani R. "An Introduction to Statistical Learning", 2nd ed., 2021, ch. 8.2.2.
- Strobl C., Boulesteix A.-L., Zeileis A., Hothorn T. "Bias in Random Forest Variable Importance Measures", BMC Bioinformatics 8, 2007.
- scikit-learn documentation, "Forests of randomized trees": https://scikit-learn.org/stable/modules/ensemble.html#forest