⚖️ AI Ethics, Society & Careers · Lecture 3 of 17

Fairness Metrics: Demographic Parity, Equalised Odds and Calibration

We formalise group fairness — demographic parity, equal opportunity, equalised odds, predictive parity and calibration — compute them with Fairlearn, and prove why several cannot hold simultaneously except in special cases.

To manage fairness, we must measure it. Researchers have proposed many formal definitions, most of which compare statistics of a model's predictions across groups defined by a protected attribute $A$. Understanding these definitions — and the tensions between them — is essential for anyone evaluating models that affect people.

Notation#

  • $Y \in \{0, 1\}$: true outcome (e.g. 1 = needs urgent support).
  • $\hat{Y} \in \{0, 1\}$: model decision; $S$: model score.
  • $A$: group membership (e.g. $a$ and $b$).

Independence: demographic (statistical) parity#

$$ P(\hat{Y} = 1 \mid A = a) = P(\hat{Y} = 1 \mid A = b) $$

Equal selection rates across groups. The disparate impact ratio $\frac{P(\hat{Y}=1 \mid A=a)}{P(\hat{Y}=1 \mid A=b)}$ is often compared with the US "four-fifths rule" (ratio below 0.8 flags possible adverse impact in employment contexts).

Appropriate when the outcome variable is itself biased or when equal allocation is a goal. Problematic when base rates genuinely differ — forcing equal selection can mean accepting less qualified candidates in one group or rejecting needy people in another.

Separation: error-rate parity#

Equal opportunity (Hardt, Price & Srebro, 2016): equal true positive rates —

$$ P(\hat{Y} = 1 \mid Y = 1, A = a) = P(\hat{Y} = 1 \mid Y = 1, A = b) $$

Those who truly need support have the same chance of being identified, whatever their group.

Equalised odds: equal true positive rates and equal false positive rates:

$$ P(\hat{Y} = 1 \mid Y = y, A = a) = P(\hat{Y} = 1 \mid Y = y, A = b), \quad y \in \{0, 1\} $$

Appropriate when labels are trustworthy and the harms are the errors themselves.

Sufficiency: predictive parity and calibration#

Predictive parity: equal precision (positive predictive value) —

$$ P(Y = 1 \mid \hat{Y} = 1, A = a) = P(Y = 1 \mid \hat{Y} = 1, A = b) $$

Calibration within groups: for every score $s$,

$$ P(Y = 1 \mid S = s, A = a) = P(Y = 1 \mid S = s, A = b) = s $$

A score of 0.7 means a 70% chance for everyone. Important when scores are communicated as risks and used by humans.

The impossibility results#

Kleinberg, Mullainathan and Raghavan (2016) and Chouldechova (2017) proved that, when base rates differ between groups ($P(Y = 1 \mid A = a) \ne P(Y = 1 \mid A = b)$), a classifier generally cannot simultaneously satisfy calibration (or predictive parity) and equal false positive and false negative rates — except in degenerate cases (perfect prediction).

Chouldechova's identity makes the tension concrete. For each group with prevalence $p$:

$$ \text{FPR} = \frac{p}{1 - p}\cdot\frac{1 - \text{PPV}}{\text{PPV}}\cdot(1 - \text{FNR}) $$

If PPV and FNR are equal across groups but prevalences $p$ differ, the FPRs must differ. This is exactly the COMPAS controversy: the tool's developers pointed to predictive parity; ProPublica pointed to unequal false positive rates. Both were measuring real properties; they could not both be equalised.

Computing fairness metrics with Fairlearn#

python
import numpy as np
from fairlearn.metrics import (MetricFrame, selection_rate, true_positive_rate, false_positive_rate,
                               demographic_parity_difference, equalized_odds_difference)
from sklearn.metrics import precision_score

rng = np.random.default_rng(1)
n = 4000
A = rng.choice(["group_a", "group_b"], n, p=[0.7, 0.3])
y = (rng.random(n) < np.where(A == "group_a", 0.30, 0.20)).astype(int)     # different base rates
score = np.clip(0.5 * y + rng.normal(0.25, 0.2, n) + np.where(A == "group_b", -0.05, 0), 0, 1)
y_pred = (score >= 0.5).astype(int)

mf = MetricFrame(metrics={"selection_rate": selection_rate, "TPR": true_positive_rate,
                          "FPR": false_positive_rate, "precision": precision_score},
                 y_true=y, y_pred=y_pred, sensitive_features=A)
print(mf.by_group.round(3))
print("demographic parity difference:", round(demographic_parity_difference(y, y_pred, sensitive_features=A), 3))
print("equalized odds difference:   ", round(equalized_odds_difference(y, y_pred, sensitive_features=A), 3))

Fairlearn also provides mitigation algorithms (ExponentiatedGradient with constraints such as EqualizedOdds, and ThresholdOptimizer); AIF360 is another comprehensive toolkit.

Beyond group metrics#

  • Individual fairness (Dwork et al., 2012): similar individuals should be treated similarly — requires a task-appropriate similarity measure.
  • Counterfactual fairness (Kusner et al., 2017): a decision should not change if the protected attribute had been different, with causally downstream features changed accordingly — requires a causal model.
  • Uncertainty: with small groups, metric differences may be noise — report confidence intervals (bootstrap).
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

⚖️ AI Ethics, Society & Careers

Bias and Fairness in Machine Learning

What does it mean for a model to be fair, and where does unfairness come from? We examine sources of bias, the failure of "fairness through unawareness", proxy variables, and practical strategies across the ML pipeline.

Intermediate⏱ 5 min#258
⚖️ AI Ethics, Society & Careers

Explainable AI: LIME, SHAP and Interpretable Models

Why did the model decide that? We distinguish interpretable models from post-hoc explanations, derive Shapley values and SHAP, explain LIME, cover global vs local explanations and counterfactuals, and discuss the limits of explanations.

Intermediate⏱ 5 min#260
⚖️ AI Ethics, Society & Careers

Why AI Ethics Matters: Principles for Responsible Engineers

AI systems make or shape decisions about people at scale. We examine real harms, the major principles of responsible AI, why good intentions are not enough, and how ethics becomes concrete engineering practice.

Beginner⏱ 5 min#257