๐Ÿ‘๏ธ Computer Vision ยท Lecture 23 of 27

OCR and Document AI: From Scanned Forms to Structured Data

Much of the world's information is locked in scanned forms, receipts and handwritten records. We cover text detection, recognition with CRNN and CTC, layout-aware models, multilingual challenges and end-to-end document understanding.

Registration forms, identity documents, invoices, historical archives, handwritten case notes โ€” enormous amounts of information exist only as images of text. Optical Character Recognition (OCR) and broader Document AI convert them into searchable, structured data. For organisations processing thousands of paper forms, this can save enormous amounts of manual data entry โ€” but errors can have real consequences for the people whose records are processed.

The classical OCR pipeline#

  1. Preprocessing: deskew, denoise, binarise (e.g. Otsu or adaptive thresholding), correct perspective (homography for phone photos).
  2. Text detection: find regions containing text โ€” lines or words.
  3. Text recognition: convert each region's image into a character string.
  4. Post-processing: language models, dictionaries and validation rules (dates, ID number checksums).
  5. Information extraction: map text to fields ("Name", "Date of birth").

Text detection#

Scene and document text varies in orientation, size and shape. Deep detectors include EAST (predicts rotated boxes per pixel), CRAFT (predicts character regions and the affinity between neighbouring characters, handling curved text) and DBNet (differentiable binarisation of a text probability map). For clean scanned documents, simpler line segmentation often suffices.

Text recognition with CRNN and CTC#

A line image has variable width and text has variable length, and we rarely know exactly which pixels correspond to which character. The CRNN architecture (Shi et al., 2015) solved this elegantly:

  1. A CNN extracts a feature sequence along the horizontal axis (one vector per column-slice).
  2. A bidirectional LSTM models context along the sequence.
  3. Connectionist Temporal Classification (CTC) loss (Graves et al., 2006) trains without character-level alignment.

CTC adds a blank symbol and sums over all alignments that collapse to the target string: repeated characters are merged, then blanks removed ("hh-e-ll-ll-oo" โ†’ "hello"; a blank between the two l's preserves the double letter). The loss is

$$ \mathcal{L}_{\text{CTC}} = -\log\sum_{\pi \in \mathcal{B}^{-1}(\mathbf{y})}\prod_{t=1}^{T}p(\pi_t \mid \mathbf{x}) $$

computed efficiently by dynamic programming (a forwardโ€“backward algorithm, like HMMs).

python
import torch
import torch.nn as nn

T, N, C = 32, 4, 37                         # time steps, batch, classes (blank + 26 letters + 10 digits)
log_probs = torch.randn(T, N, C).log_softmax(2).requires_grad_()   # model output per column
targets = torch.randint(1, C, (N, 8))        # label indices (0 is reserved for blank)
loss = nn.CTCLoss(blank=0)(log_probs, targets,
                           input_lengths=torch.full((N,), T), target_lengths=torch.full((N,), 8))
loss.backward(); print(loss.item())

def greedy_decode(log_probs_1seq, alphabet):
    best = log_probs_1seq.argmax(-1).tolist()
    out, prev = [], None
    for k in best:
        if k != prev and k != 0:
            out.append(alphabet[k - 1])
        prev = k
    return "".join(out)

Transformer-based recognisers (e.g. TrOCR, pairing a ViT encoder with a text decoder) now achieve excellent results, including on handwriting.

Practical OCR tools#

python
# pip install pytesseract easyocr   (Tesseract binary must also be installed)
import easyocr
reader = easyocr.Reader(["en", "bn"])        # English and Bangla
for box, text, conf in reader.readtext("form_photo.jpg"):
    print(f"{conf:.2f}  {text}")

Tesseract, EasyOCR, PaddleOCR and cloud APIs differ in language coverage, accuracy on handwriting and deployment options (on-premise vs cloud โ€” important for sensitive documents).

Layout and document understanding#

Documents are not just text: meaning depends on layout โ€” tables, keyโ€“value pairs, headers, checkboxes. Modern Document AI models combine three signals:

  • Text (OCR tokens),
  • Layout (2-D positions of tokens),
  • Image (visual appearance).

The LayoutLM family adds 2-D position embeddings to BERT-like models and pretrains on millions of document pages; it excels at form understanding (extracting fields) and document classification. OCR-free models such as Donut read document images directly and output structured JSON. Large multimodal language models can now answer questions about documents and extract fields from prompts โ€” powerful, but they must be validated carefully because they can hallucinate plausible values.

Multilingual and handwritten challenges#

  • Scripts with complex shaping and conjuncts (Bangla, Devanagari, Arabic with right-to-left text and cursive joining) are harder and often under-resourced.
  • Handwriting varies enormously; historical documents add degradation.
  • Mixed-language forms are common in humanitarian and administrative settings.

Fine-tuning recognisers on in-domain, local-script samples usually yields large gains.

Evaluating OCR#

  • Character Error Rate (CER) and Word Error Rate (WER): edit distance (substitutions + deletions + insertions) divided by reference length.
  • Field-level accuracy for extraction: exact match per field matters more than average CER โ€” one wrong digit in an ID number invalidates the field.
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST ยท Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

๐Ÿ‘๏ธ Computer Vision

3-D Vision: Stereo, Depth Estimation, Point Clouds and NeRF

Images are 2-D projections of a 3-D world. We cover camera geometry, stereo and monocular depth, structure from motion, point-cloud networks like PointNet, and neural scene representations such as NeRF and Gaussian splatting.

Advancedโฑ 5 min#156
๐Ÿ‘๏ธ Computer Vision

AI in Medical Imaging: Opportunities, Pitfalls and Validation

Deep learning can detect disease in X-rays, retinal scans and pathology slides. We survey modalities and tasks, discuss data and labelling challenges, shortcut learning, rigorous clinical validation, and deployment responsibilities.

Intermediateโฑ 5 min#158
๐Ÿ‘๏ธ Computer Vision

Video Understanding: Action Recognition and Temporal Modelling

Video adds time to vision. We study optical flow, two-stream networks, 3-D convolutions, (2+1)-D factorisation, video transformers and tracking, plus the computational tricks that make video models practical.

Advancedโฑ 5 min#155