Friedman's gradient boosting is elegant but, in its original form, slow on large datasets. Starting in 2014, three open-source libraries transformed it into the most successful method for tabular data: XGBoost, LightGBM and CatBoost. Knowing what each innovated — and how to use them without overfitting — is essential for any practising data scientist.
XGBoost: regularised, second-order boosting#
XGBoost (Chen & Guestrin, 2016) introduced several ideas:
1. A regularised objective. For $M$ trees $f_m$ with $T$ leaves and leaf weights $\mathbf{w}$:
2. Second-order (Newton) approximation. Expanding the loss to second order around the current prediction with gradients $g_i$ and Hessians $h_i$, the optimal weight of leaf $j$ containing examples $I_j$ is
and the gain from a split into left and right children is
Splits with negative gain are pruned — $\gamma$ acts as a minimum gain requirement. Using curvature (Hessians) makes each step more accurate than first-order gradient boosting.
3. Engineering: sparsity-aware splits that learn a default direction for missing values, weighted quantile sketches for candidate splits, cache-aware blocks, out-of-core computation and distributed training.
LightGBM: speed at scale#
LightGBM (Microsoft, 2017) focused on speed and memory:
- Histogram-based splitting — bucket continuous features into (e.g.) 255 bins, so finding splits costs $O(\text{bins})$ instead of $O(n)$.
- Leaf-wise (best-first) growth — grow the leaf with the largest gain anywhere in the tree, instead of level by level. It reaches lower loss with fewer leaves, but can overfit on small data — control it with
num_leavesandmin_data_in_leaf. - GOSS (Gradient-based One-Side Sampling) — keep all examples with large gradients and randomly sample those with small gradients, reweighting to stay unbiased.
- EFB (Exclusive Feature Bundling) — merge sparse features that are rarely non-zero together.
- Native categorical splits.
CatBoost: categorical features done right#
CatBoost (Yandex, 2017–2018) targeted two problems:
- Categorical features. Target encoding (replacing a category with the mean target) leaks the label if computed on the same data. CatBoost uses ordered target statistics: process examples in a random permutation, and encode each example using only the targets of examples before it.
- Prediction shift. Standard boosting computes residuals using a model trained on the same examples, biasing gradients. Ordered boosting uses separate models so each example's residual comes from a model that did not see it.
- Symmetric (oblivious) trees — the same split is used across an entire level. They are fast at inference and act as regularisation.
CatBoost often performs well with default settings, especially on data with many categorical columns.
Comparison#
| XGBoost | LightGBM | CatBoost | |
|---|---|---|---|
| Split finding | Exact or histogram | Histogram | Histogram |
| Tree growth | Level-wise (default) | Leaf-wise | Symmetric (oblivious) |
| Categorical handling | Native (recent versions) or encode | Native | Ordered target statistics (best-in-class) |
| Speed on large data | Fast | Fastest (typically) | Fast; slower to train on some data |
| Defaults | Good | Need care on small data | Very good |
| GPU support | Yes | Yes | Yes |
A practical template#
import numpy as np
import lightgbm as lgb
import xgboost as xgb
from catboost import CatBoostClassifier
from sklearn.datasets import fetch_openml
from sklearn.model_selection import train_test_split
from sklearn.metrics import roc_auc_score
df = fetch_openml("adult", version=2, as_frame=True).frame
y = (df.pop("class") == ">50K").astype(int)
cat_cols = df.select_dtypes("category").columns.tolist()
X_tr, X_va, y_tr, y_va = train_test_split(df, y, test_size=0.2, stratify=y, random_state=0)
# LightGBM
lgbm = lgb.LGBMClassifier(n_estimators=5000, learning_rate=0.03, num_leaves=31,
min_child_samples=40, subsample=0.8, subsample_freq=1,
colsample_bytree=0.8, reg_lambda=1.0, verbose=-1)
lgbm.fit(X_tr, y_tr, eval_set=[(X_va, y_va)], callbacks=[lgb.early_stopping(200, verbose=False)])
print("LightGBM AUC:", round(roc_auc_score(y_va, lgbm.predict_proba(X_va)[:, 1]), 4))
# XGBoost (native categorical support)
xgbm = xgb.XGBClassifier(n_estimators=5000, learning_rate=0.03, max_depth=6, subsample=0.8,
colsample_bytree=0.8, reg_lambda=1.0, tree_method="hist",
enable_categorical=True, early_stopping_rounds=200, eval_metric="auc")
xgbm.fit(X_tr, y_tr, eval_set=[(X_va, y_va)], verbose=False)
print("XGBoost AUC:", round(roc_auc_score(y_va, xgbm.predict_proba(X_va)[:, 1]), 4))
# CatBoost
cb = CatBoostClassifier(iterations=5000, learning_rate=0.05, depth=6, eval_metric="AUC",
od_type="Iter", od_wait=200, verbose=False)
cb.fit(X_tr.astype({c: str for c in cat_cols}), y_tr, cat_features=cat_cols,
eval_set=(X_va.astype({c: str for c in cat_cols}), y_va))
print("CatBoost AUC:", round(roc_auc_score(y_va, cb.predict_proba(X_va.astype({c: str for c in cat_cols}))[:, 1]), 4))Tuning priorities#
learning_rate+ number of trees via early stopping.- Tree complexity:
max_depth(XGBoost/CatBoost) ornum_leaves(LightGBM), plusmin_child_samples/min_child_weight. - Row and column subsampling (0.6–0.9).
- L1/L2 regularisation on leaf weights.
- For imbalanced data:
scale_pos_weightor class weights, and evaluate with PR-AUC.
Use a Bayesian optimiser such as Optuna for efficient search. Beyond that, gains usually come from features, not hyperparameters.
Useful extras#
- Monotonic constraints — force predictions to increase with a feature (e.g. risk with age) for sensible, auditable models.
- Custom objectives — supply gradient and Hessian functions.
- SHAP values — fast exact explanations for tree ensembles (TreeSHAP), built into all three libraries.