In October 2018, Google researchers released BERT (Devlin et al.). Fine-tuned with a single extra layer, it set new state-of-the-art results on eleven NLP benchmarks at once, including question answering and natural-language inference. BERT established the pretrain-then-fine-tune paradigm that still dominates NLP, and encoder models descended from it remain the workhorses for classification, extraction and embedding tasks in industry.
The key idea: deep bidirectionality#
Earlier pretrained language models were unidirectional: GPT-1 read left-to-right; ELMo concatenated separate left-to-right and right-to-left LSTMs. But understanding a word often requires context on both sides: "The bank of the river flooded" vs "The bank approved the loan". A standard left-to-right language-model objective cannot use right context, and naively conditioning on both sides would let each word "see itself" indirectly through the layers.
BERT's solution: masked language modelling (MLM).
Pretraining objectives#
1. Masked Language Modelling. Randomly select 15% of input tokens. Of these:
- 80% are replaced with
[MASK]; - 10% are replaced with a random token;
- 10% are left unchanged.
The model predicts the original tokens at the selected positions using the full bidirectional context. The 80/10/10 mix reduces the mismatch between pretraining (which sees [MASK]) and fine-tuning (which never does), and forces the model to keep good representations of every token.
2. Next Sentence Prediction (NSP). Given sentence pairs A and B, predict whether B actually follows A (50%) or is random (50%). Intended to help tasks involving sentence pairs โ later work (RoBERTa) found NSP unnecessary or even harmful.
Pretraining data: BooksCorpus and English Wikipedia (about 3.3 billion words). BERT-Base: 12 layers, hidden size 768, 12 heads, 110M parameters. BERT-Large: 24 layers, 1024 hidden, 16 heads, 340M parameters.
Input representation#
[CLS] the clinic opens at nine [SEP] where is it located ? [SEP]Each token's input = token embedding (WordPiece, 30,000 vocabulary) + segment embedding (A or B) + learned position embedding. The final hidden state of [CLS] serves as an aggregate representation for classification.
Fine-tuning patterns#
| Task | Input | Output layer |
|---|---|---|
| Single-sentence classification (sentiment) | [CLS] text [SEP] | Linear on [CLS] |
| Sentence-pair classification (entailment, paraphrase) | [CLS] A [SEP] B [SEP] | Linear on [CLS] |
| Token classification (NER) | [CLS] text [SEP] | Linear on every token |
| Extractive QA (SQuAD) | [CLS] question [SEP] passage [SEP] | Start/end logits over passage tokens |
All parameters are fine-tuned for a few epochs with a small learning rate (2e-5 to 5e-5).
from transformers import pipeline, AutoTokenizer, AutoModelForMaskedLM
import torch
fill = pipeline("fill-mask", model="bert-base-uncased")
for r in fill("The doctor told the patient to take the [MASK] twice a day.")[:3]:
print(round(r["score"], 3), r["token_str"])
tok = AutoTokenizer.from_pretrained("bert-base-uncased")
enc = tok("The clinic opens at nine.", "Where is it located?")
print(tok.convert_ids_to_tokens(enc["input_ids"]))
print(enc["token_type_ids"]) # segment ids: 0 for sentence A, 1 for sentence BWhy BERT worked#
- Transfer: general linguistic knowledge from billions of words transfers to tasks with few labels.
- Bidirectional context at every layer.
- Minimal task-specific architecture โ one pretrained model, many tasks.
Probing studies found BERT's layers encode a rough hierarchy: surface features in lower layers, syntax in middle layers, and more semantic, task-relevant information in upper layers.
Successors and improvements#
- RoBERTa (2019): same architecture, better training โ more data (160 GB), longer training, larger batches, dynamic masking, no NSP. Significantly better, showing BERT was undertrained.
- ALBERT (2019): parameter sharing across layers and factorised embeddings for smaller models.
- DistilBERT (2019): knowledge-distilled, 40% smaller and 60% faster while retaining most of BERT's accuracy.
- ELECTRA (2020): replaced token detection โ a small generator replaces some tokens and the main model classifies every token as original or replaced. Learning from all tokens (not only 15%) makes pretraining much more sample-efficient.
- DeBERTa (2020โ): disentangled attention (separate content and relative-position vectors) and an enhanced mask decoder; strong results on understanding benchmarks.
- SpanBERT: masks contiguous spans โ better for QA and coreference.
- Multilingual: mBERT (104 languages) and XLM-RoBERTa (100 languages, 2.5 TB of filtered web data) enable cross-lingual transfer โ fine-tune on English labels and apply to other languages. Language-specific models exist for many languages, including BanglaBERT for Bangla.
- Domain-specific: BioBERT, ClinicalBERT, SciBERT, LegalBERT, FinBERT โ continued pretraining on domain text.
- ModernBERT (2024): updated encoder with RoPE, longer context and efficient attention.
Limitations#
- Not generative: MLM encoders cannot naturally generate fluent long text.
- Context length: 512 tokens for original BERT.
- Pretrainโfine-tune mismatch from
[MASK]tokens. - Biases from pretraining text.