⚖️ AI Ethics, Society & Careers · Lecture 4 of 17

Explainable AI: LIME, SHAP and Interpretable Models

Why did the model decide that? We distinguish interpretable models from post-hoc explanations, derive Shapley values and SHAP, explain LIME, cover global vs local explanations and counterfactuals, and discuss the limits of explanations.

A caseworker is told that a household's priority score is low. A doctor sees a model flag an X-ray. A loan applicant is rejected. Each reasonably asks: why? Explanations help people trust (or appropriately distrust) models, detect errors and bias, comply with regulations that require meaningful information about automated decisions, and contest outcomes. Explainable AI (XAI) provides tools for this — with important limitations.

Interpretable by design vs post-hoc explanation#

  • Intrinsically interpretable models: linear/logistic regression with few features, small decision trees, rule lists, generalised additive models (GAMs, e.g. Explainable Boosting Machines). Their structure is the explanation.
  • Post-hoc explanations: methods that explain a trained black-box model (gradient boosting, neural networks) from the outside.

Cynthia Rudin (2019) argued forcefully that for high-stakes decisions, we should use interpretable models whenever they perform comparably — which, for many tabular problems, they do — rather than explaining black boxes with approximations that may be unfaithful.

Global vs local explanations#

  • Global: how does the model behave overall? (Which features matter most? What is the shape of the relationship?)
  • Local: why did the model make this prediction for this case?

Global tools: permutation importance, partial dependence plots (PDP), accumulated local effects (ALE), global SHAP summaries. Local tools: SHAP values, LIME, counterfactual explanations.

Shapley values and SHAP#

From cooperative game theory (Shapley, 1953): how should a "payout" (the prediction) be fairly divided among "players" (the features)? The Shapley value of feature $i$ is its average marginal contribution over all possible orders of adding features:

$$ \phi_i = \sum_{S \subseteq F \setminus \{i\}}\frac{|S|!\,(|F| - |S| - 1)!}{|F|!}\Big[v(S \cup \{i\}) - v(S)\Big] $$

where $v(S)$ is the model's expected output when only features in $S$ are known. Shapley values uniquely satisfy desirable axioms, including efficiency (local accuracy):

$$ f(\mathbf{x}) = \phi_0 + \sum_{i=1}^{M}\phi_i $$

— the prediction equals the baseline (average prediction) plus the sum of feature contributions.

SHAP (Lundberg & Lee, 2017) made Shapley values practical: TreeSHAP computes exact values for tree ensembles in polynomial time; KernelSHAP approximates them for any model; DeepSHAP and GradientSHAP for neural networks.

python
import shap
from sklearn.datasets import fetch_california_housing
from sklearn.ensemble import HistGradientBoostingRegressor

X, y = fetch_california_housing(return_X_y=True, as_frame=True)
model = HistGradientBoostingRegressor(random_state=0).fit(X, y)

explainer = shap.Explainer(model, X.sample(200, random_state=0))   # background data defines the baseline
sv = explainer(X.iloc[:500])
shap.plots.beeswarm(sv)                   # global: feature impact across many predictions
shap.plots.waterfall(sv[0])               # local: why this particular prediction?
print("baseline + contributions = prediction:",
      round(sv[0].base_values + sv[0].values.sum(), 4), round(model.predict(X.iloc[[0]])[0], 4))

LIME#

LIME (Ribeiro, Singh & Guestrin, 2016) explains a single prediction by fitting a simple, interpretable surrogate model (e.g. sparse linear regression) on perturbed samples around the instance, weighted by proximity. For text, it removes words; for images, it hides superpixels. It is model-agnostic and intuitive, but explanations can be unstable (different runs or kernel widths give different explanations) and depend on how "neighbourhood" is defined.

Counterfactual explanations#

"Your application would have been prioritised if your household had two more dependants" — a counterfactual explanation (Wachter, Mittelstadt & Russell, 2017) states the smallest change that would alter the outcome. They are actionable and intuitive, but must respect feasibility (you cannot change your age) and can reveal how to game a system; tools like DiCE generate diverse, constrained counterfactuals.

Limits and pitfalls#

Explanations for people affected#

Legal frameworks (e.g. GDPR provisions on automated decision-making) and good practice call for meaningful information about the logic involved and ways to contest decisions. Effective explanations for affected people are short, in plain language (and their own language), focused on the main reasons and on what they can do — with a clear route to a human reviewer.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

⚖️ AI Ethics, Society & Careers

AI Safety and Alignment: Making Capable Systems Do What We Intend

As AI systems grow more capable, ensuring they pursue intended goals becomes critical. We cover specification gaming, reward hacking, goal misgeneralisation, current alignment techniques, interpretability, evaluations and governance of frontier models.

Intermediate⏱ 6 min#263
⚖️ AI Ethics, Society & Careers

Fairness Metrics: Demographic Parity, Equalised Odds and Calibration

We formalise group fairness — demographic parity, equal opportunity, equalised odds, predictive parity and calibration — compute them with Fairlearn, and prove why several cannot hold simultaneously except in special cases.

Advanced⏱ 5 min#259
⚖️ AI Ethics, Society & Careers

Privacy in Machine Learning and Differential Privacy

Models can leak the data they were trained on. We examine re-identification and model attacks, why anonymisation often fails, the mathematics of differential privacy, DP-SGD, and practical privacy-by-design.

Advanced⏱ 6 min#261