A caseworker is told that a household's priority score is low. A doctor sees a model flag an X-ray. A loan applicant is rejected. Each reasonably asks: why? Explanations help people trust (or appropriately distrust) models, detect errors and bias, comply with regulations that require meaningful information about automated decisions, and contest outcomes. Explainable AI (XAI) provides tools for this — with important limitations.
Interpretable by design vs post-hoc explanation#
- Intrinsically interpretable models: linear/logistic regression with few features, small decision trees, rule lists, generalised additive models (GAMs, e.g. Explainable Boosting Machines). Their structure is the explanation.
- Post-hoc explanations: methods that explain a trained black-box model (gradient boosting, neural networks) from the outside.
Cynthia Rudin (2019) argued forcefully that for high-stakes decisions, we should use interpretable models whenever they perform comparably — which, for many tabular problems, they do — rather than explaining black boxes with approximations that may be unfaithful.
Global vs local explanations#
- Global: how does the model behave overall? (Which features matter most? What is the shape of the relationship?)
- Local: why did the model make this prediction for this case?
Global tools: permutation importance, partial dependence plots (PDP), accumulated local effects (ALE), global SHAP summaries. Local tools: SHAP values, LIME, counterfactual explanations.
Shapley values and SHAP#
From cooperative game theory (Shapley, 1953): how should a "payout" (the prediction) be fairly divided among "players" (the features)? The Shapley value of feature $i$ is its average marginal contribution over all possible orders of adding features:
where $v(S)$ is the model's expected output when only features in $S$ are known. Shapley values uniquely satisfy desirable axioms, including efficiency (local accuracy):
— the prediction equals the baseline (average prediction) plus the sum of feature contributions.
SHAP (Lundberg & Lee, 2017) made Shapley values practical: TreeSHAP computes exact values for tree ensembles in polynomial time; KernelSHAP approximates them for any model; DeepSHAP and GradientSHAP for neural networks.
import shap
from sklearn.datasets import fetch_california_housing
from sklearn.ensemble import HistGradientBoostingRegressor
X, y = fetch_california_housing(return_X_y=True, as_frame=True)
model = HistGradientBoostingRegressor(random_state=0).fit(X, y)
explainer = shap.Explainer(model, X.sample(200, random_state=0)) # background data defines the baseline
sv = explainer(X.iloc[:500])
shap.plots.beeswarm(sv) # global: feature impact across many predictions
shap.plots.waterfall(sv[0]) # local: why this particular prediction?
print("baseline + contributions = prediction:",
round(sv[0].base_values + sv[0].values.sum(), 4), round(model.predict(X.iloc[[0]])[0], 4))LIME#
LIME (Ribeiro, Singh & Guestrin, 2016) explains a single prediction by fitting a simple, interpretable surrogate model (e.g. sparse linear regression) on perturbed samples around the instance, weighted by proximity. For text, it removes words; for images, it hides superpixels. It is model-agnostic and intuitive, but explanations can be unstable (different runs or kernel widths give different explanations) and depend on how "neighbourhood" is defined.
Counterfactual explanations#
"Your application would have been prioritised if your household had two more dependants" — a counterfactual explanation (Wachter, Mittelstadt & Russell, 2017) states the smallest change that would alter the outcome. They are actionable and intuitive, but must respect feasibility (you cannot change your age) and can reveal how to game a system; tools like DiCE generate diverse, constrained counterfactuals.
Limits and pitfalls#
Explanations for people affected#
Legal frameworks (e.g. GDPR provisions on automated decision-making) and good practice call for meaningful information about the logic involved and ways to contest decisions. Effective explanations for affected people are short, in plain language (and their own language), focused on the main reasons and on what they can do — with a clear route to a human reviewer.