Welcome to the NLP & Transformers track. Language is humanity's most powerful tool: we use it to share knowledge, coordinate, negotiate, teach and comfort. Natural Language Processing (NLP) aims to give machines the ability to understand and generate it. In the last decade, NLP went from a specialised field to the engine of AI's most visible systems. Let us begin by understanding why the problem is hard.
Why language is hard#
Ambiguity at every level:
- Lexical: "bank" (river or money), "bat" (animal or cricket equipment).
- Syntactic: "I saw the man with the telescope" — who has the telescope?
- Semantic: "Every student read a book" — the same book or different ones?
- Pragmatic: "Can you pass the salt?" is a request, not a question about ability.
- Referential: "The trophy didn't fit in the suitcase because it was too big."
Other challenges:
- Context and world knowledge — understanding requires facts not stated in the text.
- Compositionality — meanings combine systematically, but idioms ("kick the bucket") break the rules.
- Variation — dialects, slang, code-mixing ("Ami office e late hobo"), spelling errors, informal social-media text.
- Long-tail vocabulary — names, new words, technical terms.
- Linguistic diversity — about 7,000 languages, most with little digital data.
Levels of linguistic analysis#
| Level | Studies | NLP task example |
|---|---|---|
| Phonetics/phonology | Sounds | Speech recognition |
| Morphology | Word structure | Stemming, lemmatisation |
| Syntax | Sentence structure | Part-of-speech tagging, parsing |
| Semantics | Meaning | Word sense disambiguation, semantic role labelling |
| Pragmatics | Meaning in context | Intent detection, dialogue |
| Discourse | Multi-sentence structure | Coreference resolution, summarisation |
Common NLP tasks#
- Classification: sentiment, topic, spam, intent, toxicity.
- Sequence labelling: part-of-speech tags, named entities.
- Structured prediction: parsing, relation extraction.
- Sequence-to-sequence: translation, summarisation, paraphrasing.
- Question answering and retrieval.
- Dialogue and conversational agents.
- Generation: stories, code, reports.
- Speech: recognition and synthesis.
Three eras of NLP#
- Rule-based (1950s–1980s): hand-written grammars and dictionaries. The 1954 Georgetown–IBM experiment translated about 60 Russian sentences and inspired over-optimistic predictions. ELIZA (1966) mimicked conversation with patterns.
- Statistical (1990s–2010s): probabilistic models learned from corpora — n-gram language models, HMM taggers, statistical machine translation, conditional random fields. Frederick Jelinek's quip captured the mood: "Every time I fire a linguist, the performance of the speech recogniser goes up."
- Neural (2013–present): word embeddings (2013), RNN sequence-to-sequence models and attention (2014–2016), transformers (2017), pretrained models like BERT and GPT (2018–), and large language models that perform many tasks from instructions.
A tiny NLP pipeline#
import re
from collections import Counter
text = """Machine learning helps organisations respond faster.
Machine translation helps people read information in their own language."""
tokens = re.findall(r"[a-zA-Z']+", text.lower())
print(tokens[:8])
print(Counter(tokens).most_common(5))
sentences = [s.strip() for s in re.split(r"(?<=[.!?])\s+", text) if s.strip()]
print(len(sentences), "sentences")Even this naive tokenisation hides decisions: what about "don't", hyphens, numbers, emojis, or scripts without spaces between words? We study these in the next lecture.
NLP and responsibility#
Language technology shapes access to information. Speakers of low-resource languages are underserved by many systems; toxic-content filters can disproportionately flag dialects; translation errors in medical, legal or asylum contexts can have serious consequences. Throughout this track we will pay attention to multilingual performance, bias and appropriate human oversight.