"Our model does not use race or gender, so it cannot be biased." This is one of the most common — and most mistaken — beliefs in applied ML. Bias can arise from data, labels, features, objectives and deployment, even when sensitive attributes are excluded. This lecture explains how unfairness arises and what we can do about it; the next lecture formalises fairness metrics.
Sources of bias in data#
- Historical bias: data faithfully records an unjust world. A model trained on past lending or hiring decisions learns past discrimination.
- Representation (sampling) bias: some groups are under-represented. Many image datasets over-represent certain regions; many speech datasets under-represent some accents and languages.
- Measurement bias: features or labels are measured differently across groups. Arrest records measure policing intensity, not just crime; "customer satisfaction" surveys may have different response rates by group.
- Label bias: human annotators' judgements (e.g. of toxicity or "professionalism") can encode stereotypes — studies found toxicity classifiers flagging text in African-American English as offensive more often.
Sources of bias in modelling#
- Objective choice: minimising overall error favours the majority group; small groups contribute little to the average loss.
- Aggregation bias: one model for groups whose feature–outcome relationships differ (e.g. a clinical risk score calibrated on one population applied to another).
- Feedback loops: predictive policing sends more patrols where more crime was recorded, which records more crime there, reinforcing the pattern.
Why "fairness through unawareness" fails#
Removing a protected attribute (gender, ethnicity, religion, disability) does not remove its influence, because other features act as proxies: postcode, name, language, school, shopping patterns, device type. A model can reconstruct protected attributes from such correlations. Worse, dropping the attribute makes it harder to measure and correct unfairness.
import numpy as np
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import cross_val_score
rng = np.random.default_rng(0)
n = 5000
group = rng.integers(0, 2, n) # protected attribute (not given to model)
neighbourhood = np.where(rng.random(n) < 0.85, group, 1 - group) # strongly correlated proxy
income = rng.normal(50 - 8 * group, 10, n) # historical inequality in income
X_without_group = np.column_stack([neighbourhood, income])
# Can the "unaware" features predict the protected attribute?
auc = cross_val_score(LogisticRegression(), X_without_group, group, cv=5, scoring="roc_auc").mean()
print(f"protected attribute recoverable from proxies with AUC = {auc:.2f}")An AUC far above 0.5 means the model can effectively "see" the protected attribute anyway.
Strategies across the pipeline#
Before training (pre-processing)#
- Collect more representative data; involve communities in data collection.
- Audit labels for bias; improve annotation guidelines and annotator diversity.
- Reweighting or resampling to balance groups; transformations that reduce dependence on protected attributes (e.g. learning fair representations).
- Examine whether the target is a biased proxy (cost vs need).
During training (in-processing)#
- Add fairness constraints or penalties to the objective (e.g. reductions approach of Agarwal et al., 2018, implemented in Fairlearn).
- Adversarial debiasing: train the model so an adversary cannot predict the protected attribute from its representation or predictions.
- Group-robust optimisation: minimise the worst-group loss (e.g. Group DRO) rather than the average.
After training (post-processing)#
- Group-specific thresholds to equalise chosen error rates (Hardt et al., 2016) — effective, but may be legally sensitive in some jurisdictions.
- Calibration per group.
At deployment#
- Monitor outcomes by group over time.
- Provide explanations, appeals and human review.
- Limit use to validated contexts.
Fairness requires context#
Intersectionality#
Bias often concentrates at intersections — e.g. older women from a minority language group — which single-attribute analyses miss (Crenshaw's concept of intersectionality, applied to AI by Buolamwini and Gebru). Evaluate intersections where sample sizes allow, and report uncertainty where they do not.