๐Ÿ’ฌ NLP & Transformers ยท Lecture 2 of 29

Text Preprocessing: Tokenisation, Normalisation, Stemming and Lemmatisation

Raw text must be converted into units a model can process. We cover normalisation, word and sentence tokenisation, stop words, stemming versus lemmatisation, and how preprocessing needs differ for classical and neural models.

Before any model sees text, we must decide what the basic units are and how to clean them. These choices look mundane but significantly affect results. A preprocessing step that helps a bag-of-words classifier can hurt a transformer; a tokeniser designed for English may fail on Bangla or Chinese. Today we learn the standard toolkit and when to use each tool.

Normalisation#

  • Unicode normalisation (NFC/NFKC) โ€” the same visible character can have several byte encodings; normalise so they compare equal. Critical for scripts with combining marks (Bangla, Devanagari, Arabic diacritics).
  • Case folding โ€” lowercasing reduces vocabulary size but loses information ("US" vs "us", "Apple" the company).
  • Removing or replacing URLs, email addresses, numbers, HTML tags โ€” replace with placeholders (<URL>, <NUM>) rather than deleting, if they carry signal.
  • Spelling normalisation for social media ("sooo gooood" โ†’ "so good") and transliteration handling for romanised text.

Tokenisation#

Tokenisation splits text into units (tokens).

  • Whitespace/rule-based word tokenisation: split on spaces and punctuation, with rules for contractions ("don't" โ†’ "do", "n't"), abbreviations ("Dr."), hyphens, numbers ("3.14"), emails and emoticons.
  • Sentence segmentation: split at ".", "?", "!", but not in "Dr." or "3.5"; Bangla uses the dari "เฅค" as a full stop.
  • Languages without spaces (Chinese, Japanese, Thai) require word-segmentation models.
  • Subword tokenisation (BPE, WordPiece, SentencePiece) โ€” used by all modern neural models; covered in a dedicated lecture.
python
import nltk
nltk.download("punkt", quiet=True); nltk.download("punkt_tab", quiet=True)
from nltk.tokenize import word_tokenize, sent_tokenize

text = "Dr. Rahman can't attend on 3.5.2025. He'll join online โ€” see https://example.org!"
print(sent_tokenize(text))
print(word_tokenize(text))

import spacy                       # python -m spacy download en_core_web_sm
nlp = spacy.load("en_core_web_sm")
doc = nlp(text)
print([(t.text, t.lemma_, t.pos_, t.is_stop) for t in doc][:10])

Stop words#

Stop words are very frequent function words ("the", "is", "of"). Removing them shrinks bag-of-words representations and can help topic modelling or keyword search.

Stemming vs lemmatisation#

Both reduce inflected words to a common base, so "running", "runs" and "ran" can be treated as related.

Stemming chops affixes with heuristic rules. The Porter stemmer (1980) turns "studies" โ†’ "studi", "university" โ†’ "univers". Fast, language-specific, and produces non-words; can over-stem ("university" and "universe" โ†’ "univers") or under-stem.

Lemmatisation uses vocabulary and morphological analysis to return the dictionary form (lemma): "studies" โ†’ "study", "better" โ†’ "good" (as adjective), "was" โ†’ "be". More accurate but slower, and it needs part-of-speech information.

WordPorter stemLemma
studiesstudistudy
runningrunrun
betterbettergood (adj.)
waswabe

For morphologically rich languages (Bangla, Turkish, Finnish, Arabic), where one root yields many surface forms, morphological analysis matters much more than in English.

Other classical steps#

  • n-grams: contiguous sequences of $n$ tokens ("machine learning" as a bigram) capture local word order.
  • Part-of-speech tagging and named-entity recognition as features.
  • Handling negation: marking tokens after "not" ("not_good") in bag-of-words models.

Classical vs neural preprocessing#

StepBag-of-words / TF-IDF modelsPretrained transformers
LowercasingOften helpfulUse the model's own convention (cased/uncased)
Stop-word removalSometimesNo
Stemming/lemmatisationOften helpfulNo โ€” subword tokeniser handles morphology
Punctuation removalSometimesNo
TokeniserWord-levelMust use the model's own tokeniser

Building a reusable preprocessing function#

python
import re, unicodedata

def normalise(text, lower=True):
    text = unicodedata.normalize("NFKC", text)
    text = re.sub(r"https?://\S+", " <URL> ", text)
    text = re.sub(r"\S+@\S+\.\S+", " <EMAIL> ", text)
    text = re.sub(r"\d+([.,]\d+)?", " <NUM> ", text)
    text = re.sub(r"(.)\1{2,}", r"\1\1", text)          # "sooo" -> "soo"
    text = re.sub(r"\s+", " ", text).strip()
    return text.lower() if lower else text

print(normalise("Sooo happy!!! Contact me at a.b@mail.com or visit https://x.org, costs 1,500"))
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST ยท Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

๐Ÿ’ฌ NLP & Transformers

Subword Tokenisation: BPE, WordPiece, Unigram and SentencePiece

Modern language models split text into subword units. We derive byte-pair encoding step by step, compare WordPiece and Unigram LM tokenisation, discuss byte-level BPE, and examine how tokenisation affects multilingual fairness and cost.

Intermediateโฑ 5 min#168
๐Ÿ’ฌ NLP & Transformers

Introduction to Natural Language Processing

We open the NLP track with the question of why human language is so hard for machines โ€” ambiguity, context, compositionality โ€” and map the tasks, eras and methods of natural language processing.

Beginnerโฑ 4 min#162
๐Ÿ’ฌ NLP & Transformers

Bag of Words and TF-IDF: Classical Text Representation

The simplest way to turn documents into vectors is to count words. We build bag-of-words and TF-IDF representations, derive the IDF formula, use cosine similarity for retrieval, and train strong linear text classifiers.

Beginnerโฑ 5 min#164