๐Ÿ’ฌ NLP & Transformers ยท Lecture 7 of 29

Subword Tokenisation: BPE, WordPiece, Unigram and SentencePiece

Modern language models split text into subword units. We derive byte-pair encoding step by step, compare WordPiece and Unigram LM tokenisation, discuss byte-level BPE, and examine how tokenisation affects multilingual fairness and cost.

Word-level vocabularies face a dilemma: a vocabulary large enough to cover rare words becomes enormous, while any fixed vocabulary still meets unknown words. Character-level models avoid unknown words but produce very long sequences. Subword tokenisation is the practical middle ground: frequent words become single tokens, while rare words are split into meaningful pieces ("unbelievable" โ†’ "un", "believ", "able"). Every modern language model uses it, and the tokeniser quietly influences cost, context length and even fairness across languages.

Byte-Pair Encoding (BPE)#

Originally a data-compression algorithm, BPE was adapted for NLP by Sennrich, Haddow and Birch (2016). Training:

  1. Start with a vocabulary of individual characters; represent every word in the corpus as a sequence of characters (with an end-of-word marker).
  2. Count all adjacent symbol pairs, weighted by word frequency.
  3. Merge the most frequent pair into a new symbol; add it to the vocabulary.
  4. Repeat until the vocabulary reaches the desired size (e.g. 32,000โ€“100,000+).

Tokenising new text applies the learned merges in order.

python
from collections import Counter

def learn_bpe(word_freqs, num_merges):
    vocab = {tuple(w) + ("</w>",): f for w, f in word_freqs.items()}
    merges = []
    for _ in range(num_merges):
        pairs = Counter()
        for sym, f in vocab.items():
            for a, b in zip(sym, sym[1:]):
                pairs[(a, b)] += f
        if not pairs:
            break
        best = max(pairs, key=pairs.get)
        merges.append(best)
        new_vocab = {}
        for sym, f in vocab.items():
            out, i = [], 0
            while i < len(sym):
                if i < len(sym) - 1 and (sym[i], sym[i + 1]) == best:
                    out.append(sym[i] + sym[i + 1]); i += 2
                else:
                    out.append(sym[i]); i += 1
            new_vocab[tuple(out)] = f
        vocab = new_vocab
    return merges, vocab

merges, vocab = learn_bpe({"low": 5, "lower": 2, "newest": 6, "widest": 3}, 10)
print(merges[:6])          # e.g. ('e','s'), ('es','t'), ('est','</w>'), ('l','o'), ...
print(list(vocab)[:4])

This classic example from the BPE paper learns merges such as "es" โ†’ "est" โ†’ "est</w>", so "newest" and "widest" share the suffix token "est</w>".

Byte-level BPE#

GPT-2 introduced byte-level BPE: start from the 256 possible bytes of UTF-8 text instead of characters. Every possible string can be encoded โ€” no unknown tokens ever, including emojis and any script โ€” while merges still create efficient tokens for common sequences. Most GPT-style models use byte-level BPE (e.g. OpenAI's tiktoken encodings).

WordPiece#

Used by BERT. Similar to BPE, but instead of merging the most frequent pair, it merges the pair that most increases the likelihood of the training data under a unigram model โ€” approximately the pair maximising

$$ \text{score}(a, b) = \frac{\text{freq}(ab)}{\text{freq}(a)\times\text{freq}(b)} $$

favouring pairs whose parts are rarely seen apart. Continuation pieces are marked with "##": "playing" โ†’ "play", "##ing".

Unigram language model tokenisation#

Kudo (2018) took the opposite approach: start with a large candidate vocabulary and iteratively remove tokens whose removal least reduces the corpus likelihood under a unigram model, $P(\mathbf{x}) = \prod_i p(x_i)$. A word can have several segmentations; the tokeniser chooses the most probable (Viterbi) or samples segmentations during training (subword regularisation), which improves robustness.

SentencePiece#

SentencePiece (Kudo & Richardson, 2018) is a toolkit implementing BPE and Unigram that treats the input as a raw stream including spaces (represented as "โ–"). It needs no language-specific pre-tokenisation โ€” essential for languages without spaces between words โ€” and makes tokenisation fully reversible. T5, LLaMA-family and many multilingual models use SentencePiece.

python
# pip install sentencepiece transformers
import sentencepiece as spm
spm.SentencePieceTrainer.train(input="corpus.txt", model_prefix="mytok",
                               vocab_size=8000, model_type="unigram", character_coverage=0.9995)
sp = spm.SentencePieceProcessor(model_file="mytok.model")
print(sp.encode("Machine learning is transforming education.", out_type=str))

from transformers import AutoTokenizer
for name in ["bert-base-uncased", "gpt2"]:
    tok = AutoTokenizer.from_pretrained(name)
    print(name, tok.tokenize("Tokenisation influences multilingual fairness!"))

Why tokenisation matters#

  1. Sequence length and cost: models have fixed context windows, and APIs charge per token. Text that splits into more tokens costs more and fits less content.
  2. Multilingual fairness: tokenisers trained mostly on English split other languages โ€” especially non-Latin scripts โ€” into many more tokens per word. Studies (e.g. Petrov et al., 2023) found the same content can require several times more tokens in some languages than in English, meaning higher cost, shorter effective context and often lower quality for those speakers.
  3. Arithmetic and spelling: numbers split inconsistently ("12345" โ†’ "123", "45") make arithmetic harder; models see tokens, not letters, which explains difficulty with tasks like counting letters in a word.
  4. Glitch tokens: rare tokens seen seldom in training can trigger strange behaviour.
AlgorithmDirectionCriterionUsed by
BPEBottom-up mergesPair frequencyGPT family (byte-level), many LLMs
WordPieceBottom-up mergesLikelihood gainBERT
Unigram LMTop-down pruningLikelihood lossT5, ALBERT, via SentencePiece
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST ยท Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

๐Ÿ’ฌ NLP & Transformers

GloVe and FastText: Global Statistics and Subword Embeddings

GloVe learns embeddings from a global co-occurrence matrix with a weighted least-squares objective; FastText represents words as bags of character n-grams, handling rare and unseen words. We compare them with word2vec and discuss evaluation.

Intermediateโฑ 5 min#167
๐Ÿ’ฌ NLP & Transformers

Text Preprocessing: Tokenisation, Normalisation, Stemming and Lemmatisation

Raw text must be converted into units a model can process. We cover normalisation, word and sentence tokenisation, stop words, stemming versus lemmatisation, and how preprocessing needs differ for classical and neural models.

Beginnerโฑ 5 min#163
๐Ÿ’ฌ NLP & Transformers

Text Classification: From Linear Models to Fine-Tuned Transformers

Text classification is the most widely deployed NLP task. We compare TF-IDF baselines, CNN and RNN classifiers, and fine-tuned transformers, and cover label design, imbalance, multilingual data and evaluation.

Beginnerโฑ 4 min#169