Before any model sees text, we must decide what the basic units are and how to clean them. These choices look mundane but significantly affect results. A preprocessing step that helps a bag-of-words classifier can hurt a transformer; a tokeniser designed for English may fail on Bangla or Chinese. Today we learn the standard toolkit and when to use each tool.
Normalisation#
- Unicode normalisation (NFC/NFKC) โ the same visible character can have several byte encodings; normalise so they compare equal. Critical for scripts with combining marks (Bangla, Devanagari, Arabic diacritics).
- Case folding โ lowercasing reduces vocabulary size but loses information ("US" vs "us", "Apple" the company).
- Removing or replacing URLs, email addresses, numbers, HTML tags โ replace with placeholders (
<URL>,<NUM>) rather than deleting, if they carry signal. - Spelling normalisation for social media ("sooo gooood" โ "so good") and transliteration handling for romanised text.
Tokenisation#
Tokenisation splits text into units (tokens).
- Whitespace/rule-based word tokenisation: split on spaces and punctuation, with rules for contractions ("don't" โ "do", "n't"), abbreviations ("Dr."), hyphens, numbers ("3.14"), emails and emoticons.
- Sentence segmentation: split at ".", "?", "!", but not in "Dr." or "3.5"; Bangla uses the dari "เฅค" as a full stop.
- Languages without spaces (Chinese, Japanese, Thai) require word-segmentation models.
- Subword tokenisation (BPE, WordPiece, SentencePiece) โ used by all modern neural models; covered in a dedicated lecture.
import nltk
nltk.download("punkt", quiet=True); nltk.download("punkt_tab", quiet=True)
from nltk.tokenize import word_tokenize, sent_tokenize
text = "Dr. Rahman can't attend on 3.5.2025. He'll join online โ see https://example.org!"
print(sent_tokenize(text))
print(word_tokenize(text))
import spacy # python -m spacy download en_core_web_sm
nlp = spacy.load("en_core_web_sm")
doc = nlp(text)
print([(t.text, t.lemma_, t.pos_, t.is_stop) for t in doc][:10])Stop words#
Stop words are very frequent function words ("the", "is", "of"). Removing them shrinks bag-of-words representations and can help topic modelling or keyword search.
Stemming vs lemmatisation#
Both reduce inflected words to a common base, so "running", "runs" and "ran" can be treated as related.
Stemming chops affixes with heuristic rules. The Porter stemmer (1980) turns "studies" โ "studi", "university" โ "univers". Fast, language-specific, and produces non-words; can over-stem ("university" and "universe" โ "univers") or under-stem.
Lemmatisation uses vocabulary and morphological analysis to return the dictionary form (lemma): "studies" โ "study", "better" โ "good" (as adjective), "was" โ "be". More accurate but slower, and it needs part-of-speech information.
| Word | Porter stem | Lemma |
|---|---|---|
| studies | studi | study |
| running | run | run |
| better | better | good (adj.) |
| was | wa | be |
For morphologically rich languages (Bangla, Turkish, Finnish, Arabic), where one root yields many surface forms, morphological analysis matters much more than in English.
Other classical steps#
- n-grams: contiguous sequences of $n$ tokens ("machine learning" as a bigram) capture local word order.
- Part-of-speech tagging and named-entity recognition as features.
- Handling negation: marking tokens after "not" ("not_good") in bag-of-words models.
Classical vs neural preprocessing#
| Step | Bag-of-words / TF-IDF models | Pretrained transformers |
|---|---|---|
| Lowercasing | Often helpful | Use the model's own convention (cased/uncased) |
| Stop-word removal | Sometimes | No |
| Stemming/lemmatisation | Often helpful | No โ subword tokeniser handles morphology |
| Punctuation removal | Sometimes | No |
| Tokeniser | Word-level | Must use the model's own tokeniser |
Building a reusable preprocessing function#
import re, unicodedata
def normalise(text, lower=True):
text = unicodedata.normalize("NFKC", text)
text = re.sub(r"https?://\S+", " <URL> ", text)
text = re.sub(r"\S+@\S+\.\S+", " <EMAIL> ", text)
text = re.sub(r"\d+([.,]\d+)?", " <NUM> ", text)
text = re.sub(r"(.)\1{2,}", r"\1\1", text) # "sooo" -> "soo"
text = re.sub(r"\s+", " ", text).strip()
return text.lower() if lower else text
print(normalise("Sooo happy!!! Contact me at a.b@mail.com or visit https://x.org, costs 1,500"))