To manage fairness, we must measure it. Researchers have proposed many formal definitions, most of which compare statistics of a model's predictions across groups defined by a protected attribute $A$. Understanding these definitions — and the tensions between them — is essential for anyone evaluating models that affect people.
Notation#
- $Y \in \{0, 1\}$: true outcome (e.g. 1 = needs urgent support).
- $\hat{Y} \in \{0, 1\}$: model decision; $S$: model score.
- $A$: group membership (e.g. $a$ and $b$).
Independence: demographic (statistical) parity#
Equal selection rates across groups. The disparate impact ratio $\frac{P(\hat{Y}=1 \mid A=a)}{P(\hat{Y}=1 \mid A=b)}$ is often compared with the US "four-fifths rule" (ratio below 0.8 flags possible adverse impact in employment contexts).
Appropriate when the outcome variable is itself biased or when equal allocation is a goal. Problematic when base rates genuinely differ — forcing equal selection can mean accepting less qualified candidates in one group or rejecting needy people in another.
Separation: error-rate parity#
Equal opportunity (Hardt, Price & Srebro, 2016): equal true positive rates —
Those who truly need support have the same chance of being identified, whatever their group.
Equalised odds: equal true positive rates and equal false positive rates:
Appropriate when labels are trustworthy and the harms are the errors themselves.
Sufficiency: predictive parity and calibration#
Predictive parity: equal precision (positive predictive value) —
Calibration within groups: for every score $s$,
A score of 0.7 means a 70% chance for everyone. Important when scores are communicated as risks and used by humans.
The impossibility results#
Kleinberg, Mullainathan and Raghavan (2016) and Chouldechova (2017) proved that, when base rates differ between groups ($P(Y = 1 \mid A = a) \ne P(Y = 1 \mid A = b)$), a classifier generally cannot simultaneously satisfy calibration (or predictive parity) and equal false positive and false negative rates — except in degenerate cases (perfect prediction).
Chouldechova's identity makes the tension concrete. For each group with prevalence $p$:
If PPV and FNR are equal across groups but prevalences $p$ differ, the FPRs must differ. This is exactly the COMPAS controversy: the tool's developers pointed to predictive parity; ProPublica pointed to unequal false positive rates. Both were measuring real properties; they could not both be equalised.
Computing fairness metrics with Fairlearn#
import numpy as np
from fairlearn.metrics import (MetricFrame, selection_rate, true_positive_rate, false_positive_rate,
demographic_parity_difference, equalized_odds_difference)
from sklearn.metrics import precision_score
rng = np.random.default_rng(1)
n = 4000
A = rng.choice(["group_a", "group_b"], n, p=[0.7, 0.3])
y = (rng.random(n) < np.where(A == "group_a", 0.30, 0.20)).astype(int) # different base rates
score = np.clip(0.5 * y + rng.normal(0.25, 0.2, n) + np.where(A == "group_b", -0.05, 0), 0, 1)
y_pred = (score >= 0.5).astype(int)
mf = MetricFrame(metrics={"selection_rate": selection_rate, "TPR": true_positive_rate,
"FPR": false_positive_rate, "precision": precision_score},
y_true=y, y_pred=y_pred, sensitive_features=A)
print(mf.by_group.round(3))
print("demographic parity difference:", round(demographic_parity_difference(y, y_pred, sensitive_features=A), 3))
print("equalized odds difference: ", round(equalized_odds_difference(y, y_pred, sensitive_features=A), 3))Fairlearn also provides mitigation algorithms (ExponentiatedGradient with constraints such as EqualizedOdds, and ThresholdOptimizer); AIF360 is another comprehensive toolkit.
Beyond group metrics#
- Individual fairness (Dwork et al., 2012): similar individuals should be treated similarly — requires a task-appropriate similarity measure.
- Counterfactual fairness (Kusner et al., 2017): a decision should not change if the protected attribute had been different, with causally downstream features changed accordingly — requires a causal model.
- Uncertainty: with small groups, metric differences may be noise — report confidence intervals (bootstrap).