๐Ÿ‘๏ธ Computer Vision ยท Lecture 24 of 27

AI in Medical Imaging: Opportunities, Pitfalls and Validation

Deep learning can detect disease in X-rays, retinal scans and pathology slides. We survey modalities and tasks, discuss data and labelling challenges, shortcut learning, rigorous clinical validation, and deployment responsibilities.

Medical imaging is one of the most promising โ€” and most demanding โ€” application areas of computer vision. Models have matched specialists on specific tasks such as detecting diabetic retinopathy in retinal photographs or finding certain findings on chest X-rays. In places with too few radiologists, AI screening could extend care to millions. But the history of medical AI also contains many models that performed brilliantly in papers and poorly in hospitals. Today we study both the potential and the discipline required.

Modalities and tasks#

ModalityTypical tasks
Chest X-rayClassification (pneumonia, tuberculosis, nodules), triage
CT / MRI (3-D volumes)Segmentation of organs and tumours, detection of haemorrhage
Retinal fundus / OCTDiabetic retinopathy grading, glaucoma screening
Digital pathology (gigapixel slides)Tumour detection, grading, cell counting
Dermatology photographsLesion classification
UltrasoundFetal biometry, cardiac function, point-of-care screening
MammographyBreast cancer screening

Architectures are familiar: CNNs and ViTs for classification, U-Net variants (including 3-D U-Net and the self-configuring nnU-Net, a very strong baseline) for segmentation, and multiple-instance learning for gigapixel pathology slides, where a slide-level label must be learned from thousands of tiles.

Data challenges#

  • Small labelled datasets โ€” expert labelling is expensive and slow.
  • Label noise and disagreement โ€” radiologists disagree; use multiple readers, adjudication, or labels from stronger reference standards (biopsy, follow-up) where possible.
  • Class imbalance โ€” disease is rarer than normal findings.
  • Privacy โ€” de-identify DICOM headers and burned-in text; follow legal and ethical approval processes.
  • Heterogeneity โ€” different scanners, protocols, populations and hospitals.

Transfer learning, self-supervised pretraining on unlabelled scans, heavy but anatomically valid augmentation, and federated learning (training across hospitals without sharing raw data) help address these.

Shortcut learning: the central danger#

Models learn whatever predicts the label in the training data โ€” including artefacts:

  • A pneumonia classifier trained on data from several hospitals learned to recognise which hospital an image came from (hospitals had different pneumonia prevalence and scanner markers), inflating accuracy (Zech et al., 2018).
  • Studies found models using chest tubes (a treatment already given) to "detect" pneumothorax, or ruler markings in dermatology photos (rulers were more common beside suspicious lesions).
  • Deep networks have been shown able to predict patient attributes such as self-reported race from X-rays in ways humans cannot, raising concerns about hidden biases in downstream predictions.

Rigorous evaluation#

  1. Split by patient, never by image โ€” images of the same patient must not appear in both training and test sets.
  2. External validation on data from other hospitals, countries and devices.
  3. Clinically meaningful metrics: sensitivity and specificity at an operating point chosen with clinicians; negative predictive value for rule-out tools; calibration.
  4. Subgroup analysis: age, sex, ethnicity, disease severity, device type.
  5. Compare with clinicians on the same cases โ€” and, better, measure clinician + AI versus clinician alone, since most systems assist rather than replace.
  6. Prospective studies and, ideally, randomised trials measuring patient outcomes.
  7. Reporting guidelines: CLAIM, TRIPOD+AI, CONSORT-AI and STARD-AI promote transparent reporting.
python
import numpy as np
from sklearn.metrics import roc_auc_score, confusion_matrix

def clinical_report(y_true, scores, threshold, groups):
    pred = (scores >= threshold).astype(int)
    tn, fp, fn, tp = confusion_matrix(y_true, pred).ravel()
    print(f"AUC {roc_auc_score(y_true, scores):.3f} | sensitivity {tp / (tp + fn):.3f} | "
          f"specificity {tn / (tn + fp):.3f} | NPV {tn / (tn + fn):.3f}")
    for g in np.unique(groups):                      # subgroup analysis
        m = groups == g
        tn, fp, fn, tp = confusion_matrix(y_true[m], pred[m], labels=[0, 1]).ravel()
        print(f"  {g:<12} n={m.sum():>4}  sensitivity {tp / max(tp + fn, 1):.3f}  specificity {tn / max(tn + fp, 1):.3f}")

Deployment and regulation#

Medical AI is typically regulated as a medical device (e.g. by the FDA in the United States or under the Medical Device Regulation in the EU). Hundreds of AI-enabled devices have received regulatory clearance, most in radiology. Post-deployment, performance must be monitored โ€” new scanners, software updates or population changes can degrade it. Integration into clinical workflow, clear communication of uncertainty and responsibility, and avoiding automation bias (clinicians over-trusting the AI) are as important as the model.

Global health perspective#

AI screening for tuberculosis on chest X-rays has been recommended by the WHO for use in certain screening settings, and diabetic retinopathy screening tools have been deployed in primary care. In low-resource settings, benefits can be greatest โ€” and so can risks if models trained on data from high-income hospitals are applied to different populations and equipment without local validation.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST ยท Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

๐Ÿ‘๏ธ Computer Vision

Semantic Segmentation: FCN, U-Net and DeepLab

Segmentation labels every pixel. We cover fully convolutional networks, the encoderโ€“decoder U-Net with skip connections, DeepLab's atrous convolutions, loss functions like Dice, and evaluation with IoU.

Intermediateโฑ 5 min#148
๐Ÿ‘๏ธ Computer Vision

OCR and Document AI: From Scanned Forms to Structured Data

Much of the world's information is locked in scanned forms, receipts and handwritten records. We cover text detection, recognition with CRNN and CTC, layout-aware models, multilingual challenges and end-to-end document understanding.

Intermediateโฑ 5 min#157
๐Ÿ‘๏ธ Computer Vision

Vision Foundation Models: Segment Anything and Promptable Vision

Vision is following language towards general-purpose foundation models. We study the Segment Anything Model's promptable design and data engine, open-vocabulary detection, and how foundation models change vision workflows.

Advancedโฑ 5 min#159