๐Ÿ’ฌ NLP & Transformers ยท Lecture 16 of 29

BERT: Bidirectional Encoder Representations from Transformers

BERT showed that a bidirectional transformer pretrained on unlabelled text could be fine-tuned to beat task-specific models across NLP. We cover masked language modelling, input format, fine-tuning patterns, and successors such as RoBERTa, DeBERTa and multilingual encoders.

In October 2018, Google researchers released BERT (Devlin et al.). Fine-tuned with a single extra layer, it set new state-of-the-art results on eleven NLP benchmarks at once, including question answering and natural-language inference. BERT established the pretrain-then-fine-tune paradigm that still dominates NLP, and encoder models descended from it remain the workhorses for classification, extraction and embedding tasks in industry.

The key idea: deep bidirectionality#

Earlier pretrained language models were unidirectional: GPT-1 read left-to-right; ELMo concatenated separate left-to-right and right-to-left LSTMs. But understanding a word often requires context on both sides: "The bank of the river flooded" vs "The bank approved the loan". A standard left-to-right language-model objective cannot use right context, and naively conditioning on both sides would let each word "see itself" indirectly through the layers.

BERT's solution: masked language modelling (MLM).

Pretraining objectives#

1. Masked Language Modelling. Randomly select 15% of input tokens. Of these:

  • 80% are replaced with [MASK];
  • 10% are replaced with a random token;
  • 10% are left unchanged.

The model predicts the original tokens at the selected positions using the full bidirectional context. The 80/10/10 mix reduces the mismatch between pretraining (which sees [MASK]) and fine-tuning (which never does), and forces the model to keep good representations of every token.

2. Next Sentence Prediction (NSP). Given sentence pairs A and B, predict whether B actually follows A (50%) or is random (50%). Intended to help tasks involving sentence pairs โ€” later work (RoBERTa) found NSP unnecessary or even harmful.

Pretraining data: BooksCorpus and English Wikipedia (about 3.3 billion words). BERT-Base: 12 layers, hidden size 768, 12 heads, 110M parameters. BERT-Large: 24 layers, 1024 hidden, 16 heads, 340M parameters.

Input representation#

text
[CLS] the clinic opens at nine [SEP] where is it located ? [SEP]

Each token's input = token embedding (WordPiece, 30,000 vocabulary) + segment embedding (A or B) + learned position embedding. The final hidden state of [CLS] serves as an aggregate representation for classification.

Fine-tuning patterns#

TaskInputOutput layer
Single-sentence classification (sentiment)[CLS] text [SEP]Linear on [CLS]
Sentence-pair classification (entailment, paraphrase)[CLS] A [SEP] B [SEP]Linear on [CLS]
Token classification (NER)[CLS] text [SEP]Linear on every token
Extractive QA (SQuAD)[CLS] question [SEP] passage [SEP]Start/end logits over passage tokens

All parameters are fine-tuned for a few epochs with a small learning rate (2e-5 to 5e-5).

python
from transformers import pipeline, AutoTokenizer, AutoModelForMaskedLM
import torch

fill = pipeline("fill-mask", model="bert-base-uncased")
for r in fill("The doctor told the patient to take the [MASK] twice a day.")[:3]:
    print(round(r["score"], 3), r["token_str"])

tok = AutoTokenizer.from_pretrained("bert-base-uncased")
enc = tok("The clinic opens at nine.", "Where is it located?")
print(tok.convert_ids_to_tokens(enc["input_ids"]))
print(enc["token_type_ids"])                 # segment ids: 0 for sentence A, 1 for sentence B

Why BERT worked#

  1. Transfer: general linguistic knowledge from billions of words transfers to tasks with few labels.
  2. Bidirectional context at every layer.
  3. Minimal task-specific architecture โ€” one pretrained model, many tasks.

Probing studies found BERT's layers encode a rough hierarchy: surface features in lower layers, syntax in middle layers, and more semantic, task-relevant information in upper layers.

Successors and improvements#

  • RoBERTa (2019): same architecture, better training โ€” more data (160 GB), longer training, larger batches, dynamic masking, no NSP. Significantly better, showing BERT was undertrained.
  • ALBERT (2019): parameter sharing across layers and factorised embeddings for smaller models.
  • DistilBERT (2019): knowledge-distilled, 40% smaller and 60% faster while retaining most of BERT's accuracy.
  • ELECTRA (2020): replaced token detection โ€” a small generator replaces some tokens and the main model classifies every token as original or replaced. Learning from all tokens (not only 15%) makes pretraining much more sample-efficient.
  • DeBERTa (2020โ€“): disentangled attention (separate content and relative-position vectors) and an enhanced mask decoder; strong results on understanding benchmarks.
  • SpanBERT: masks contiguous spans โ€” better for QA and coreference.
  • Multilingual: mBERT (104 languages) and XLM-RoBERTa (100 languages, 2.5 TB of filtered web data) enable cross-lingual transfer โ€” fine-tune on English labels and apply to other languages. Language-specific models exist for many languages, including BanglaBERT for Bangla.
  • Domain-specific: BioBERT, ClinicalBERT, SciBERT, LegalBERT, FinBERT โ€” continued pretraining on domain text.
  • ModernBERT (2024): updated encoder with RoPE, longer context and efficient attention.

Limitations#

  • Not generative: MLM encoders cannot naturally generate fluent long text.
  • Context length: 512 tokens for original BERT.
  • Pretrainโ€“fine-tune mismatch from [MASK] tokens.
  • Biases from pretraining text.
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST ยท Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

๐Ÿ’ฌ NLP & Transformers

Text Classification: From Linear Models to Fine-Tuned Transformers

Text classification is the most widely deployed NLP task. We compare TF-IDF baselines, CNN and RNN classifiers, and fine-tuned transformers, and cover label design, imbalance, multilingual data and evaluation.

Beginnerโฑ 4 min#169
๐Ÿ’ฌ NLP & Transformers

Fine-Tuning Pretrained Language Models: A Practical Guide

How to adapt a pretrained language model to your task reliably โ€” choosing a model, preparing data, hyperparameters, handling small data and instability, continued pretraining, and evaluating properly.

Intermediateโฑ 5 min#180
๐Ÿ’ฌ NLP & Transformers

Positional Encodings: Sinusoidal, Learned, RoPE and ALiBi

Attention is order-blind, so transformers need positional information. We compare absolute sinusoidal and learned encodings with relative methods โ€” rotary embeddings (RoPE) and ALiBi โ€” and discuss extending context length.

Advancedโฑ 5 min#176