📈 Machine Learning · Lecture 22 of 47

XGBoost, LightGBM and CatBoost: Modern Gradient Boosting Libraries

Three libraries turned gradient boosting into an industrial tool. We compare their key innovations — second-order optimisation, histogram splits, leaf-wise growth, GOSS, ordered target statistics — and show how to use each well.

Friedman's gradient boosting is elegant but, in its original form, slow on large datasets. Starting in 2014, three open-source libraries transformed it into the most successful method for tabular data: XGBoost, LightGBM and CatBoost. Knowing what each innovated — and how to use them without overfitting — is essential for any practising data scientist.

XGBoost: regularised, second-order boosting#

XGBoost (Chen & Guestrin, 2016) introduced several ideas:

1. A regularised objective. For $M$ trees $f_m$ with $T$ leaves and leaf weights $\mathbf{w}$:

$$ \mathcal{L} = \sum_i L(y_i, \hat{y}_i) + \sum_m\Omega(f_m), \qquad \Omega(f) = \gamma T + \frac{1}{2}\lambda\|\mathbf{w}\|^2 $$

2. Second-order (Newton) approximation. Expanding the loss to second order around the current prediction with gradients $g_i$ and Hessians $h_i$, the optimal weight of leaf $j$ containing examples $I_j$ is

$$ w_j^* = -\frac{\sum_{i \in I_j}g_i}{\sum_{i \in I_j}h_i + \lambda} $$

and the gain from a split into left and right children is

$$ \text{Gain} = \frac{1}{2}\left[\frac{G_L^2}{H_L + \lambda} + \frac{G_R^2}{H_R + \lambda} - \frac{(G_L + G_R)^2}{H_L + H_R + \lambda}\right] - \gamma $$

Splits with negative gain are pruned — $\gamma$ acts as a minimum gain requirement. Using curvature (Hessians) makes each step more accurate than first-order gradient boosting.

3. Engineering: sparsity-aware splits that learn a default direction for missing values, weighted quantile sketches for candidate splits, cache-aware blocks, out-of-core computation and distributed training.

LightGBM: speed at scale#

LightGBM (Microsoft, 2017) focused on speed and memory:

  • Histogram-based splitting — bucket continuous features into (e.g.) 255 bins, so finding splits costs $O(\text{bins})$ instead of $O(n)$.
  • Leaf-wise (best-first) growth — grow the leaf with the largest gain anywhere in the tree, instead of level by level. It reaches lower loss with fewer leaves, but can overfit on small data — control it with num_leaves and min_data_in_leaf.
  • GOSS (Gradient-based One-Side Sampling) — keep all examples with large gradients and randomly sample those with small gradients, reweighting to stay unbiased.
  • EFB (Exclusive Feature Bundling) — merge sparse features that are rarely non-zero together.
  • Native categorical splits.

CatBoost: categorical features done right#

CatBoost (Yandex, 2017–2018) targeted two problems:

  • Categorical features. Target encoding (replacing a category with the mean target) leaks the label if computed on the same data. CatBoost uses ordered target statistics: process examples in a random permutation, and encode each example using only the targets of examples before it.
  • Prediction shift. Standard boosting computes residuals using a model trained on the same examples, biasing gradients. Ordered boosting uses separate models so each example's residual comes from a model that did not see it.
  • Symmetric (oblivious) trees — the same split is used across an entire level. They are fast at inference and act as regularisation.

CatBoost often performs well with default settings, especially on data with many categorical columns.

Comparison#

XGBoostLightGBMCatBoost
Split findingExact or histogramHistogramHistogram
Tree growthLevel-wise (default)Leaf-wiseSymmetric (oblivious)
Categorical handlingNative (recent versions) or encodeNativeOrdered target statistics (best-in-class)
Speed on large dataFastFastest (typically)Fast; slower to train on some data
DefaultsGoodNeed care on small dataVery good
GPU supportYesYesYes

A practical template#

python
import numpy as np
import lightgbm as lgb
import xgboost as xgb
from catboost import CatBoostClassifier
from sklearn.datasets import fetch_openml
from sklearn.model_selection import train_test_split
from sklearn.metrics import roc_auc_score

df = fetch_openml("adult", version=2, as_frame=True).frame
y = (df.pop("class") == ">50K").astype(int)
cat_cols = df.select_dtypes("category").columns.tolist()
X_tr, X_va, y_tr, y_va = train_test_split(df, y, test_size=0.2, stratify=y, random_state=0)

# LightGBM
lgbm = lgb.LGBMClassifier(n_estimators=5000, learning_rate=0.03, num_leaves=31,
                          min_child_samples=40, subsample=0.8, subsample_freq=1,
                          colsample_bytree=0.8, reg_lambda=1.0, verbose=-1)
lgbm.fit(X_tr, y_tr, eval_set=[(X_va, y_va)], callbacks=[lgb.early_stopping(200, verbose=False)])
print("LightGBM AUC:", round(roc_auc_score(y_va, lgbm.predict_proba(X_va)[:, 1]), 4))

# XGBoost (native categorical support)
xgbm = xgb.XGBClassifier(n_estimators=5000, learning_rate=0.03, max_depth=6, subsample=0.8,
                         colsample_bytree=0.8, reg_lambda=1.0, tree_method="hist",
                         enable_categorical=True, early_stopping_rounds=200, eval_metric="auc")
xgbm.fit(X_tr, y_tr, eval_set=[(X_va, y_va)], verbose=False)
print("XGBoost AUC:", round(roc_auc_score(y_va, xgbm.predict_proba(X_va)[:, 1]), 4))

# CatBoost
cb = CatBoostClassifier(iterations=5000, learning_rate=0.05, depth=6, eval_metric="AUC",
                        od_type="Iter", od_wait=200, verbose=False)
cb.fit(X_tr.astype({c: str for c in cat_cols}), y_tr, cat_features=cat_cols,
       eval_set=(X_va.astype({c: str for c in cat_cols}), y_va))
print("CatBoost AUC:", round(roc_auc_score(y_va, cb.predict_proba(X_va.astype({c: str for c in cat_cols}))[:, 1]), 4))

Tuning priorities#

  1. learning_rate + number of trees via early stopping.
  2. Tree complexity: max_depth (XGBoost/CatBoost) or num_leaves (LightGBM), plus min_child_samples/min_child_weight.
  3. Row and column subsampling (0.6–0.9).
  4. L1/L2 regularisation on leaf weights.
  5. For imbalanced data: scale_pos_weight or class weights, and evaluate with PR-AUC.

Use a Bayesian optimiser such as Optuna for efficient search. Beyond that, gains usually come from features, not hyperparameters.

Useful extras#

  • Monotonic constraints — force predictions to increase with a feature (e.g. risk with age) for sensible, auditable models.
  • Custom objectives — supply gradient and Hessian functions.
  • SHAP values — fast exact explanations for tree ensembles (TreeSHAP), built into all three libraries.
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

📈 Machine Learning

Boosting II: Gradient Boosting Machines

Gradient boosting performs gradient descent in function space, fitting each new tree to the negative gradient of the loss. We derive the algorithm, explain shrinkage and subsampling, and tune it properly.

Advanced⏱ 5 min#070
📈 Machine Learning

Support Vector Machines: Maximum-Margin Classification

Among all separating hyperplanes, SVMs pick the one with the widest margin. We derive the hard- and soft-margin formulations, the hinge loss, support vectors and the role of the C parameter.

Intermediate⏱ 5 min#072
📈 Machine Learning

Boosting I: AdaBoost and the Power of Weak Learners

Can many weak rules of thumb combine into a strong classifier? AdaBoost answered yes. We walk through the reweighting algorithm, derive it as exponential-loss minimisation, and discuss its margins and sensitivity to noise.

Intermediate⏱ 5 min#069