Machine translation (MT) was one of the first proposed applications of computers, and today it lets billions of people read content in languages they do not speak. For displaced people navigating services in a new country, good translation can be the difference between understanding one's rights and not. This lecture follows MT from its statistical roots to modern neural and massively multilingual systems — and examines where it still fails.
A brief history#
- Rule-based MT (1950s–1980s): dictionaries and hand-written grammar transfer rules. Brittle and expensive to build.
- Statistical MT (SMT) (1990s–2016): IBM alignment models and phrase-based systems learned translation probabilities from parallel corpora, combined in a log-linear model with an n-gram language model. Google Translate used phrase-based SMT for years.
- Neural MT (NMT) (2014–): sequence-to-sequence models with attention (Bahdanau et al., 2015) quickly surpassed SMT. In 2016, Google's GNMT reported large reductions in translation errors for major language pairs.
- Transformer NMT (2017–): the Transformer was introduced for translation and became the standard.
- Massively multilingual and LLM-based MT (2019–): single models translating between hundreds of languages; large language models translating via prompting.
The NMT model#
An encoder reads the source sentence; a decoder generates the target autoregressively, attending to the encoder states:
Training maximises the likelihood of reference translations (cross-entropy with teacher forcing, often with label smoothing). Decoding uses beam search with length normalisation.
Data#
- Parallel corpora: sentence-aligned translations — parliamentary proceedings, UN documents, subtitles, religious texts, web-mined pairs.
- Quality filtering matters: web-mined data contains misalignments and machine-translated text.
- Shared subword vocabularies (BPE/SentencePiece) across source and target languages let rare words and names be copied or transliterated.
Back-translation#
Monolingual target-language text is far more abundant than parallel text. Back-translation (Sennrich et al., 2016): train a reverse model (target → source), translate monolingual target sentences into synthetic sources, and add the synthetic pairs to training. The decoder learns from real, fluent target text. It is one of the most effective techniques, especially for low-resource pairs.
Multilingual NMT#
A single model can translate among many languages by prepending a target-language tag (e.g. <2bn>). Benefits: transfer from high-resource to related low-resource languages, and even zero-shot translation between pairs never seen together. Meta's NLLB-200 ("No Language Left Behind", 2022) supports 200 languages, trained with mined data and evaluated on the FLORES-200 benchmark covering many under-served languages.
from transformers import pipeline
translator = pipeline("translation", model="facebook/nllb-200-distilled-600M",
src_lang="eng_Latn", tgt_lang="ben_Beng")
print(translator("Registration is free. Please bring an identity document if you have one.",
max_length=100)[0]["translation_text"])Evaluation#
BLEU (Papineni et al., 2002) measures n-gram overlap between the output and reference translations:
where $p_n$ is modified n-gram precision and BP penalises translations shorter than the reference ($r$ = reference length, $c$ = candidate length). BLEU is cheap and reproducible (use sacreBLEU for standardised scores) but correlates imperfectly with human judgement: valid paraphrases score poorly.
chrF uses character n-grams — better for morphologically rich languages. Learned metrics such as COMET and BLEURT, trained to predict human quality judgements from embeddings, correlate substantially better with humans and are now standard alongside BLEU. Human evaluation (e.g. direct assessment, MQM error annotation) remains the gold standard.
import sacrebleu
refs = [["The clinic opens at nine in the morning."]]
hyp = ["The clinic opens at 9 am."]
print(sacrebleu.corpus_bleu(hyp, refs).score, sacrebleu.corpus_chrf(hyp, refs).score)Persistent challenges#
- Low-resource languages: little parallel data, poor tokenisation, few evaluation sets.
- Domain shift: medical, legal and humanitarian terminology.
- Hallucinations: fluent output unrelated to the source, especially under domain shift or with noisy input.
- Gender and bias: translating from gender-neutral languages (e.g. Bangla's "সে", Turkish "o") into English often defaults to stereotypes ("he is a doctor, she is a nurse").
- Formality, culture and idioms.
- Document-level context: pronouns and terminology consistency across sentences.