If you can use only one algorithm for a new tabular dataset with no time to tune, the random forest is an excellent choice. Introduced by Leo Breiman in 2001, it combines bagging with an extra dose of randomness, producing models that are accurate, robust to noise and outliers, and nearly impossible to badly misconfigure.
The key idea: decorrelate the trees#
Bagged trees are correlated because they all tend to pick the same strong features near the root. Recall the ensemble variance formula:
To reduce $\rho$, random forests restrict each split to a random subset of $m$ features (out of $d$). Strong features are unavailable at some splits, forcing trees to explore other structure. Each tree becomes slightly worse individually, but the ensemble improves because errors are less correlated.
The algorithm#
For $b = 1, \dots, B$:
- Draw a bootstrap sample of the training data.
- Grow a tree on it. At each node:
- select $m$ features at random;
- find the best split among only those $m$ features;
- split, and recurse โ typically growing trees deep, with no pruning.
Predict by averaging (regression) or majority vote/averaged probabilities (classification).
Typical defaults: $m = \sqrt{d}$ for classification, $m = d/3$ (or all features, in some libraries) for regression.
Hyperparameters that matter#
| Hyperparameter | Effect | Guidance |
|---|---|---|
n_estimators ($B$) | More trees โ lower variance; never overfits | 200โ1000; stop when OOB error plateaus |
max_features ($m$) | Smaller โ more decorrelation, weaker trees | Tune among $\sqrt{d}$, $\log_2 d$, 0.3โ0.5ยท$d$ |
min_samples_leaf | Larger โ smoother, less variance | 1โ10; larger for noisy regression |
max_depth | Limit tree depth | Usually unlimited; limit for speed/memory |
class_weight | Reweight classes | "balanced" for imbalanced data |
Random forests are famous for working well with defaults โ a big practical advantage.
import numpy as np
from sklearn.datasets import fetch_california_housing
from sklearn.model_selection import train_test_split
from sklearn.ensemble import RandomForestRegressor
from sklearn.tree import DecisionTreeRegressor
from sklearn.metrics import mean_absolute_error
X, y = fetch_california_housing(return_X_y=True, as_frame=True)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.2, random_state=0)
tree = DecisionTreeRegressor(random_state=0).fit(X_tr, y_tr)
print("single tree MAE:", round(mean_absolute_error(y_te, tree.predict(X_te)), 3))
for m in [1.0, 0.5, 0.33]:
rf = RandomForestRegressor(n_estimators=300, max_features=m, oob_score=True,
n_jobs=-1, random_state=0).fit(X_tr, y_tr)
print(f"RF max_features={m}: test MAE {mean_absolute_error(y_te, rf.predict(X_te)):.3f}"
f" OOB R^2 {rf.oob_score_:.3f}")(With max_features=1.0, the forest is plain bagging.)
Feature importance โ handle with care#
Random forests offer two importance measures:
- Mean decrease in impurity (MDI) โ fast, computed during training, but biased towards continuous and high-cardinality features and computed on training data.
- Permutation importance โ shuffle one feature in validation data and measure the drop in performance. More reliable, but can understate the importance of correlated features (the model uses the correlated partner instead).
from sklearn.inspection import permutation_importance
rf = RandomForestRegressor(n_estimators=300, n_jobs=-1, random_state=0).fit(X_tr, y_tr)
pi = permutation_importance(rf, X_te, y_te, n_repeats=10, random_state=0, n_jobs=-1)
for i in pi.importances_mean.argsort()[::-1]:
print(f"{X.columns[i]:<12} {pi.importances_mean[i]:.3f} ยฑ {pi.importances_std[i]:.3f}")Other useful by-products#
- OOB error โ free validation.
- Proximity matrix โ how often two examples land in the same leaf; useful for clustering, outlier detection and imputing missing values.
- Quantile regression forests โ keep all leaf targets to estimate prediction intervals.
- Isolation Forest โ a related randomised-tree method for anomaly detection.
Strengths and weaknesses#
Strengths: strong accuracy with little tuning; robust to outliers and irrelevant features; handles non-linearity and interactions; parallelisable; no feature scaling; OOB estimates.
Weaknesses: less interpretable than a single tree; large memory footprint and slower prediction than a linear model; cannot extrapolate beyond the training target range; usually slightly less accurate than well-tuned gradient boosting on tabular data.
Random forest or gradient boosting?#
| Consideration | Random forest | Gradient boosting |
|---|---|---|
| Tuning effort | Low | Moderate to high |
| Peak accuracy on tabular data | Very good | Often best |
| Overfitting by adding trees | No | Yes (needs early stopping) |
| Training parallelism | Across trees | Within trees |
| Noisy labels | Very robust | Can chase noise |