Word-level vocabularies face a dilemma: a vocabulary large enough to cover rare words becomes enormous, while any fixed vocabulary still meets unknown words. Character-level models avoid unknown words but produce very long sequences. Subword tokenisation is the practical middle ground: frequent words become single tokens, while rare words are split into meaningful pieces ("unbelievable" โ "un", "believ", "able"). Every modern language model uses it, and the tokeniser quietly influences cost, context length and even fairness across languages.
Byte-Pair Encoding (BPE)#
Originally a data-compression algorithm, BPE was adapted for NLP by Sennrich, Haddow and Birch (2016). Training:
- Start with a vocabulary of individual characters; represent every word in the corpus as a sequence of characters (with an end-of-word marker).
- Count all adjacent symbol pairs, weighted by word frequency.
- Merge the most frequent pair into a new symbol; add it to the vocabulary.
- Repeat until the vocabulary reaches the desired size (e.g. 32,000โ100,000+).
Tokenising new text applies the learned merges in order.
from collections import Counter
def learn_bpe(word_freqs, num_merges):
vocab = {tuple(w) + ("</w>",): f for w, f in word_freqs.items()}
merges = []
for _ in range(num_merges):
pairs = Counter()
for sym, f in vocab.items():
for a, b in zip(sym, sym[1:]):
pairs[(a, b)] += f
if not pairs:
break
best = max(pairs, key=pairs.get)
merges.append(best)
new_vocab = {}
for sym, f in vocab.items():
out, i = [], 0
while i < len(sym):
if i < len(sym) - 1 and (sym[i], sym[i + 1]) == best:
out.append(sym[i] + sym[i + 1]); i += 2
else:
out.append(sym[i]); i += 1
new_vocab[tuple(out)] = f
vocab = new_vocab
return merges, vocab
merges, vocab = learn_bpe({"low": 5, "lower": 2, "newest": 6, "widest": 3}, 10)
print(merges[:6]) # e.g. ('e','s'), ('es','t'), ('est','</w>'), ('l','o'), ...
print(list(vocab)[:4])This classic example from the BPE paper learns merges such as "es" โ "est" โ "est</w>", so "newest" and "widest" share the suffix token "est</w>".
Byte-level BPE#
GPT-2 introduced byte-level BPE: start from the 256 possible bytes of UTF-8 text instead of characters. Every possible string can be encoded โ no unknown tokens ever, including emojis and any script โ while merges still create efficient tokens for common sequences. Most GPT-style models use byte-level BPE (e.g. OpenAI's tiktoken encodings).
WordPiece#
Used by BERT. Similar to BPE, but instead of merging the most frequent pair, it merges the pair that most increases the likelihood of the training data under a unigram model โ approximately the pair maximising
favouring pairs whose parts are rarely seen apart. Continuation pieces are marked with "##": "playing" โ "play", "##ing".
Unigram language model tokenisation#
Kudo (2018) took the opposite approach: start with a large candidate vocabulary and iteratively remove tokens whose removal least reduces the corpus likelihood under a unigram model, $P(\mathbf{x}) = \prod_i p(x_i)$. A word can have several segmentations; the tokeniser chooses the most probable (Viterbi) or samples segmentations during training (subword regularisation), which improves robustness.
SentencePiece#
SentencePiece (Kudo & Richardson, 2018) is a toolkit implementing BPE and Unigram that treats the input as a raw stream including spaces (represented as "โ"). It needs no language-specific pre-tokenisation โ essential for languages without spaces between words โ and makes tokenisation fully reversible. T5, LLaMA-family and many multilingual models use SentencePiece.
# pip install sentencepiece transformers
import sentencepiece as spm
spm.SentencePieceTrainer.train(input="corpus.txt", model_prefix="mytok",
vocab_size=8000, model_type="unigram", character_coverage=0.9995)
sp = spm.SentencePieceProcessor(model_file="mytok.model")
print(sp.encode("Machine learning is transforming education.", out_type=str))
from transformers import AutoTokenizer
for name in ["bert-base-uncased", "gpt2"]:
tok = AutoTokenizer.from_pretrained(name)
print(name, tok.tokenize("Tokenisation influences multilingual fairness!"))Why tokenisation matters#
- Sequence length and cost: models have fixed context windows, and APIs charge per token. Text that splits into more tokens costs more and fits less content.
- Multilingual fairness: tokenisers trained mostly on English split other languages โ especially non-Latin scripts โ into many more tokens per word. Studies (e.g. Petrov et al., 2023) found the same content can require several times more tokens in some languages than in English, meaning higher cost, shorter effective context and often lower quality for those speakers.
- Arithmetic and spelling: numbers split inconsistently ("12345" โ "123", "45") make arithmetic harder; models see tokens, not letters, which explains difficulty with tasks like counting letters in a word.
- Glitch tokens: rare tokens seen seldom in training can trigger strange behaviour.
| Algorithm | Direction | Criterion | Used by |
|---|---|---|---|
| BPE | Bottom-up merges | Pair frequency | GPT family (byte-level), many LLMs |
| WordPiece | Bottom-up merges | Likelihood gain | BERT |
| Unigram LM | Top-down pruning | Likelihood loss | T5, ALBERT, via SentencePiece |