Medical imaging is one of the most promising โ and most demanding โ application areas of computer vision. Models have matched specialists on specific tasks such as detecting diabetic retinopathy in retinal photographs or finding certain findings on chest X-rays. In places with too few radiologists, AI screening could extend care to millions. But the history of medical AI also contains many models that performed brilliantly in papers and poorly in hospitals. Today we study both the potential and the discipline required.
Modalities and tasks#
| Modality | Typical tasks |
|---|---|
| Chest X-ray | Classification (pneumonia, tuberculosis, nodules), triage |
| CT / MRI (3-D volumes) | Segmentation of organs and tumours, detection of haemorrhage |
| Retinal fundus / OCT | Diabetic retinopathy grading, glaucoma screening |
| Digital pathology (gigapixel slides) | Tumour detection, grading, cell counting |
| Dermatology photographs | Lesion classification |
| Ultrasound | Fetal biometry, cardiac function, point-of-care screening |
| Mammography | Breast cancer screening |
Architectures are familiar: CNNs and ViTs for classification, U-Net variants (including 3-D U-Net and the self-configuring nnU-Net, a very strong baseline) for segmentation, and multiple-instance learning for gigapixel pathology slides, where a slide-level label must be learned from thousands of tiles.
Data challenges#
- Small labelled datasets โ expert labelling is expensive and slow.
- Label noise and disagreement โ radiologists disagree; use multiple readers, adjudication, or labels from stronger reference standards (biopsy, follow-up) where possible.
- Class imbalance โ disease is rarer than normal findings.
- Privacy โ de-identify DICOM headers and burned-in text; follow legal and ethical approval processes.
- Heterogeneity โ different scanners, protocols, populations and hospitals.
Transfer learning, self-supervised pretraining on unlabelled scans, heavy but anatomically valid augmentation, and federated learning (training across hospitals without sharing raw data) help address these.
Shortcut learning: the central danger#
Models learn whatever predicts the label in the training data โ including artefacts:
- A pneumonia classifier trained on data from several hospitals learned to recognise which hospital an image came from (hospitals had different pneumonia prevalence and scanner markers), inflating accuracy (Zech et al., 2018).
- Studies found models using chest tubes (a treatment already given) to "detect" pneumothorax, or ruler markings in dermatology photos (rulers were more common beside suspicious lesions).
- Deep networks have been shown able to predict patient attributes such as self-reported race from X-rays in ways humans cannot, raising concerns about hidden biases in downstream predictions.
Rigorous evaluation#
- Split by patient, never by image โ images of the same patient must not appear in both training and test sets.
- External validation on data from other hospitals, countries and devices.
- Clinically meaningful metrics: sensitivity and specificity at an operating point chosen with clinicians; negative predictive value for rule-out tools; calibration.
- Subgroup analysis: age, sex, ethnicity, disease severity, device type.
- Compare with clinicians on the same cases โ and, better, measure clinician + AI versus clinician alone, since most systems assist rather than replace.
- Prospective studies and, ideally, randomised trials measuring patient outcomes.
- Reporting guidelines: CLAIM, TRIPOD+AI, CONSORT-AI and STARD-AI promote transparent reporting.
import numpy as np
from sklearn.metrics import roc_auc_score, confusion_matrix
def clinical_report(y_true, scores, threshold, groups):
pred = (scores >= threshold).astype(int)
tn, fp, fn, tp = confusion_matrix(y_true, pred).ravel()
print(f"AUC {roc_auc_score(y_true, scores):.3f} | sensitivity {tp / (tp + fn):.3f} | "
f"specificity {tn / (tn + fp):.3f} | NPV {tn / (tn + fn):.3f}")
for g in np.unique(groups): # subgroup analysis
m = groups == g
tn, fp, fn, tp = confusion_matrix(y_true[m], pred[m], labels=[0, 1]).ravel()
print(f" {g:<12} n={m.sum():>4} sensitivity {tp / max(tp + fn, 1):.3f} specificity {tn / max(tn + fp, 1):.3f}")Deployment and regulation#
Medical AI is typically regulated as a medical device (e.g. by the FDA in the United States or under the Medical Device Regulation in the EU). Hundreds of AI-enabled devices have received regulatory clearance, most in radiology. Post-deployment, performance must be monitored โ new scanners, software updates or population changes can degrade it. Integration into clinical workflow, clear communication of uncertainty and responsibility, and avoiding automation bias (clinicians over-trusting the AI) are as important as the model.
Global health perspective#
AI screening for tuberculosis on chest X-rays has been recommended by the WHO for use in certain screening settings, and diabetic retinopathy screening tools have been deployed in primary care. In low-resource settings, benefits can be greatest โ and so can risks if models trained on data from high-income hospitals are applied to different populations and equipment without local validation.