💬 NLP & Transformers · Lecture 1 of 29

Introduction to Natural Language Processing

We open the NLP track with the question of why human language is so hard for machines — ambiguity, context, compositionality — and map the tasks, eras and methods of natural language processing.

Welcome to the NLP & Transformers track. Language is humanity's most powerful tool: we use it to share knowledge, coordinate, negotiate, teach and comfort. Natural Language Processing (NLP) aims to give machines the ability to understand and generate it. In the last decade, NLP went from a specialised field to the engine of AI's most visible systems. Let us begin by understanding why the problem is hard.

Why language is hard#

Ambiguity at every level:

  • Lexical: "bank" (river or money), "bat" (animal or cricket equipment).
  • Syntactic: "I saw the man with the telescope" — who has the telescope?
  • Semantic: "Every student read a book" — the same book or different ones?
  • Pragmatic: "Can you pass the salt?" is a request, not a question about ability.
  • Referential: "The trophy didn't fit in the suitcase because it was too big."

Other challenges:

  • Context and world knowledge — understanding requires facts not stated in the text.
  • Compositionality — meanings combine systematically, but idioms ("kick the bucket") break the rules.
  • Variation — dialects, slang, code-mixing ("Ami office e late hobo"), spelling errors, informal social-media text.
  • Long-tail vocabulary — names, new words, technical terms.
  • Linguistic diversity — about 7,000 languages, most with little digital data.

Levels of linguistic analysis#

LevelStudiesNLP task example
Phonetics/phonologySoundsSpeech recognition
MorphologyWord structureStemming, lemmatisation
SyntaxSentence structurePart-of-speech tagging, parsing
SemanticsMeaningWord sense disambiguation, semantic role labelling
PragmaticsMeaning in contextIntent detection, dialogue
DiscourseMulti-sentence structureCoreference resolution, summarisation

Common NLP tasks#

  • Classification: sentiment, topic, spam, intent, toxicity.
  • Sequence labelling: part-of-speech tags, named entities.
  • Structured prediction: parsing, relation extraction.
  • Sequence-to-sequence: translation, summarisation, paraphrasing.
  • Question answering and retrieval.
  • Dialogue and conversational agents.
  • Generation: stories, code, reports.
  • Speech: recognition and synthesis.

Three eras of NLP#

  1. Rule-based (1950s–1980s): hand-written grammars and dictionaries. The 1954 Georgetown–IBM experiment translated about 60 Russian sentences and inspired over-optimistic predictions. ELIZA (1966) mimicked conversation with patterns.
  2. Statistical (1990s–2010s): probabilistic models learned from corpora — n-gram language models, HMM taggers, statistical machine translation, conditional random fields. Frederick Jelinek's quip captured the mood: "Every time I fire a linguist, the performance of the speech recogniser goes up."
  3. Neural (2013–present): word embeddings (2013), RNN sequence-to-sequence models and attention (2014–2016), transformers (2017), pretrained models like BERT and GPT (2018–), and large language models that perform many tasks from instructions.

A tiny NLP pipeline#

python
import re
from collections import Counter

text = """Machine learning helps organisations respond faster.
Machine translation helps people read information in their own language."""

tokens = re.findall(r"[a-zA-Z']+", text.lower())
print(tokens[:8])
print(Counter(tokens).most_common(5))

sentences = [s.strip() for s in re.split(r"(?<=[.!?])\s+", text) if s.strip()]
print(len(sentences), "sentences")

Even this naive tokenisation hides decisions: what about "don't", hyphens, numbers, emojis, or scripts without spaces between words? We study these in the next lecture.

NLP and responsibility#

Language technology shapes access to information. Speakers of low-resource languages are underserved by many systems; toxic-content filters can disproportionately flag dialects; translation errors in medical, legal or asylum contexts can have serious consequences. Throughout this track we will pay attention to multilingual performance, bias and appropriate human oversight.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

💬 NLP & Transformers

Text Preprocessing: Tokenisation, Normalisation, Stemming and Lemmatisation

Raw text must be converted into units a model can process. We cover normalisation, word and sentence tokenisation, stop words, stemming versus lemmatisation, and how preprocessing needs differ for classical and neural models.

Beginner⏱ 5 min#163
💬 NLP & Transformers

Bag of Words and TF-IDF: Classical Text Representation

The simplest way to turn documents into vectors is to count words. We build bag-of-words and TF-IDF representations, derive the IDF formula, use cosine similarity for retrieval, and train strong linear text classifiers.

Beginner⏱ 5 min#164
💬 NLP & Transformers

N-gram Language Models, Smoothing and Perplexity

A language model assigns probabilities to sequences of words. We derive n-gram models from the chain rule and Markov assumption, fix zero probabilities with smoothing, generate text, and evaluate with perplexity.

Intermediate⏱ 5 min#165