Registration forms, identity documents, invoices, historical archives, handwritten case notes โ enormous amounts of information exist only as images of text. Optical Character Recognition (OCR) and broader Document AI convert them into searchable, structured data. For organisations processing thousands of paper forms, this can save enormous amounts of manual data entry โ but errors can have real consequences for the people whose records are processed.
The classical OCR pipeline#
- Preprocessing: deskew, denoise, binarise (e.g. Otsu or adaptive thresholding), correct perspective (homography for phone photos).
- Text detection: find regions containing text โ lines or words.
- Text recognition: convert each region's image into a character string.
- Post-processing: language models, dictionaries and validation rules (dates, ID number checksums).
- Information extraction: map text to fields ("Name", "Date of birth").
Text detection#
Scene and document text varies in orientation, size and shape. Deep detectors include EAST (predicts rotated boxes per pixel), CRAFT (predicts character regions and the affinity between neighbouring characters, handling curved text) and DBNet (differentiable binarisation of a text probability map). For clean scanned documents, simpler line segmentation often suffices.
Text recognition with CRNN and CTC#
A line image has variable width and text has variable length, and we rarely know exactly which pixels correspond to which character. The CRNN architecture (Shi et al., 2015) solved this elegantly:
- A CNN extracts a feature sequence along the horizontal axis (one vector per column-slice).
- A bidirectional LSTM models context along the sequence.
- Connectionist Temporal Classification (CTC) loss (Graves et al., 2006) trains without character-level alignment.
CTC adds a blank symbol and sums over all alignments that collapse to the target string: repeated characters are merged, then blanks removed ("hh-e-ll-ll-oo" โ "hello"; a blank between the two l's preserves the double letter). The loss is
computed efficiently by dynamic programming (a forwardโbackward algorithm, like HMMs).
import torch
import torch.nn as nn
T, N, C = 32, 4, 37 # time steps, batch, classes (blank + 26 letters + 10 digits)
log_probs = torch.randn(T, N, C).log_softmax(2).requires_grad_() # model output per column
targets = torch.randint(1, C, (N, 8)) # label indices (0 is reserved for blank)
loss = nn.CTCLoss(blank=0)(log_probs, targets,
input_lengths=torch.full((N,), T), target_lengths=torch.full((N,), 8))
loss.backward(); print(loss.item())
def greedy_decode(log_probs_1seq, alphabet):
best = log_probs_1seq.argmax(-1).tolist()
out, prev = [], None
for k in best:
if k != prev and k != 0:
out.append(alphabet[k - 1])
prev = k
return "".join(out)Transformer-based recognisers (e.g. TrOCR, pairing a ViT encoder with a text decoder) now achieve excellent results, including on handwriting.
Practical OCR tools#
# pip install pytesseract easyocr (Tesseract binary must also be installed)
import easyocr
reader = easyocr.Reader(["en", "bn"]) # English and Bangla
for box, text, conf in reader.readtext("form_photo.jpg"):
print(f"{conf:.2f} {text}")Tesseract, EasyOCR, PaddleOCR and cloud APIs differ in language coverage, accuracy on handwriting and deployment options (on-premise vs cloud โ important for sensitive documents).
Layout and document understanding#
Documents are not just text: meaning depends on layout โ tables, keyโvalue pairs, headers, checkboxes. Modern Document AI models combine three signals:
- Text (OCR tokens),
- Layout (2-D positions of tokens),
- Image (visual appearance).
The LayoutLM family adds 2-D position embeddings to BERT-like models and pretrains on millions of document pages; it excels at form understanding (extracting fields) and document classification. OCR-free models such as Donut read document images directly and output structured JSON. Large multimodal language models can now answer questions about documents and extract fields from prompts โ powerful, but they must be validated carefully because they can hallucinate plausible values.
Multilingual and handwritten challenges#
- Scripts with complex shaping and conjuncts (Bangla, Devanagari, Arabic with right-to-left text and cursive joining) are harder and often under-resourced.
- Handwriting varies enormously; historical documents add degradation.
- Mixed-language forms are common in humanitarian and administrative settings.
Fine-tuning recognisers on in-domain, local-script samples usually yields large gains.
Evaluating OCR#
- Character Error Rate (CER) and Word Error Rate (WER): edit distance (substitutions + deletions + insertions) divided by reference length.
- Field-level accuracy for extraction: exact match per field matters more than average CER โ one wrong digit in an ID number invalidates the field.