💬

NLP & Transformers

Language models, embeddings, attention, BERT, GPT, speech and multilingual NLP.

  1. 01Introduction to Natural Language ProcessingWe open the NLP track with the question of why human language is so hard for machines — ambiguity, context, compositionality — and map the tasks, eras and methods of natural language processing.Beginner4 min
  2. 02Text Preprocessing: Tokenisation, Normalisation, Stemming and LemmatisationRaw text must be converted into units a model can process. We cover normalisation, word and sentence tokenisation, stop words, stemming versus lemmatisation, and how preprocessing needs differ for classical and neural models.Beginner5 min
  3. 03Bag of Words and TF-IDF: Classical Text RepresentationThe simplest way to turn documents into vectors is to count words. We build bag-of-words and TF-IDF representations, derive the IDF formula, use cosine similarity for retrieval, and train strong linear text classifiers.Beginner5 min
  4. 04N-gram Language Models, Smoothing and PerplexityA language model assigns probabilities to sequences of words. We derive n-gram models from the chain rule and Markov assumption, fix zero probabilities with smoothing, generate text, and evaluate with perplexity.Intermediate5 min
  5. 05Word2Vec: Learning Word Embeddings from ContextWord2Vec learns dense word vectors by predicting context words. We derive the skip-gram and CBOW objectives, negative sampling, and explore analogies, similarity and the limitations of static embeddings.Intermediate5 min
  6. 06GloVe and FastText: Global Statistics and Subword EmbeddingsGloVe learns embeddings from a global co-occurrence matrix with a weighted least-squares objective; FastText represents words as bags of character n-grams, handling rare and unseen words. We compare them with word2vec and discuss evaluation.Intermediate5 min
  7. 07Subword Tokenisation: BPE, WordPiece, Unigram and SentencePieceModern language models split text into subword units. We derive byte-pair encoding step by step, compare WordPiece and Unigram LM tokenisation, discuss byte-level BPE, and examine how tokenisation affects multilingual fairness and cost.Intermediate5 min
  8. 08Text Classification: From Linear Models to Fine-Tuned TransformersText classification is the most widely deployed NLP task. We compare TF-IDF baselines, CNN and RNN classifiers, and fine-tuned transformers, and cover label design, imbalance, multilingual data and evaluation.Beginner4 min
  9. 09Sentiment Analysis: Opinions, Aspects and NuanceSentiment analysis detects opinions and emotions in text. We compare lexicon-based and learned approaches, tackle negation and sarcasm, extend to aspect-based sentiment and emotions, and discuss evaluation and misuse.Beginner5 min
  10. 10Named Entity Recognition: Finding People, Places and OrganisationsNER locates and classifies entity mentions in text. We cover BIO tagging, feature-based and neural approaches, transformer token classification with subword alignment, entity-level evaluation, and domain adaptation.Intermediate5 min
  11. 11Part-of-Speech Tagging and Conditional Random FieldsPOS tagging assigns grammatical categories to words. We compare HMM taggers with discriminative Conditional Random Fields, derive the CRF likelihood and Viterbi decoding, and see why CRF layers still appear on top of neural encoders.Advanced5 min
  12. 12Neural Machine Translation: From Seq2Seq to Multilingual TransformersMachine translation is one of NLP's oldest and most impactful tasks. We trace its evolution to neural systems, cover training data, subword vocabularies, back-translation, evaluation with BLEU and COMET, and the challenges of low-resource languages.Intermediate5 min
  13. 13The Transformer Architecture Explained, Block by BlockThe 2017 Transformer replaced recurrence with attention and became the foundation of modern AI. We walk through embeddings, positional encoding, multi-head self-attention, feed-forward layers, residuals, normalisation, masking and the encoder–decoder design.Intermediate6 min
  14. 14Self-Attention in Depth: Intuition, Complexity and VariantsA deeper look at self-attention — what attention heads learn, the geometry of queries and keys, computational complexity, causal masking, KV caching, and efficient variants such as multi-query and grouped-query attention.Advanced6 min
  15. 15Positional Encodings: Sinusoidal, Learned, RoPE and ALiBiAttention is order-blind, so transformers need positional information. We compare absolute sinusoidal and learned encodings with relative methods — rotary embeddings (RoPE) and ALiBi — and discuss extending context length.Advanced5 min
  16. 16BERT: Bidirectional Encoder Representations from TransformersBERT showed that a bidirectional transformer pretrained on unlabelled text could be fine-tuned to beat task-specific models across NLP. We cover masked language modelling, input format, fine-tuning patterns, and successors such as RoBERTa, DeBERTa and multilingual encoders.Intermediate5 min
  17. 17The GPT Family: Autoregressive Language Models from GPT-1 to TodayGPT models are decoder-only transformers trained to predict the next token. We trace GPT-1 through GPT-3's few-shot learning to instruction-tuned assistants, and explain why next-token prediction at scale produces broad capabilities.Intermediate5 min
  18. 18T5 and BART: Encoder–Decoder Pretraining and Text-to-Text LearningT5 casts every NLP task as text in, text out; BART pretrains as a denoising autoencoder. We cover span corruption, the text-to-text framework, the lessons of T5's systematic study, and when encoder–decoders are the right choice.Intermediate5 min
  19. 19Fine-Tuning Pretrained Language Models: A Practical GuideHow to adapt a pretrained language model to your task reliably — choosing a model, preparing data, hyperparameters, handling small data and instability, continued pretraining, and evaluating properly.Intermediate5 min
  20. 20Question Answering: Extractive, Open-Domain and GenerativeQuestion answering systems return answers, not documents. We cover extractive reading comprehension with span prediction, open-domain retriever–reader pipelines, generative QA, evaluation metrics and the problem of unanswerable questions.Intermediate5 min
  21. 21Text Summarisation: Extractive and Abstractive MethodsSummarisation condenses documents while preserving key information. We compare extractive methods (TextRank) with abstractive neural models, evaluate with ROUGE and factual-consistency checks, and discuss long documents and hallucination.Intermediate5 min
  22. 22Evaluating NLP Systems: Perplexity, BLEU, ROUGE, BERTScore and Human JudgementHow do we know if a language system is good? We survey intrinsic and extrinsic evaluation, overlap metrics, embedding-based metrics, learned metrics, LLM judges, human evaluation, benchmarks and their pitfalls.Intermediate5 min
  23. 23Information Retrieval and Semantic SearchSearch is the most used NLP application. We cover indexing, BM25, dense bi-encoder retrieval, cross-encoder re-ranking, hybrid search, approximate nearest neighbours, and evaluation with recall@k, MRR and nDCG.Intermediate5 min
  24. 24Topic Modelling: LDA, NMF and Neural Topic ModelsTopic models discover themes in large document collections without labels. We derive Latent Dirichlet Allocation's generative story, compare it with NMF and embedding-based BERTopic, and discuss evaluating and interpreting topics.Intermediate5 min
  25. 25Multilingual and Low-Resource NLP (with a Focus on Bangla)Most of the world's languages have little digital data. We examine why this matters, how multilingual models enable cross-lingual transfer, the specific challenges of languages like Bangla, and practical strategies for building NLP in low-resource settings.Intermediate5 min
  26. 26Dialogue Systems and Chatbots: From Rules to LLM AssistantsWe compare rule-based, task-oriented and open-domain dialogue systems; cover intent detection, slot filling and dialogue state tracking; and show how LLM-based assistants with retrieval and tools are designed, evaluated and deployed safely.Intermediate5 min
  27. 27Automatic Speech Recognition: From HMMs to WhisperSpeech recognition converts audio into text. We cover audio features and spectrograms, the classical HMM–GMM pipeline, end-to-end neural models with CTC and attention, self-supervised wav2vec 2.0, Whisper, and evaluation with word error rate.Intermediate5 min
  28. 28Text-to-Speech: From Concatenation to Neural VoicesText-to-speech systems turn written text into natural-sounding speech. We cover the TTS pipeline, text normalisation and phonemes, acoustic models like Tacotron and FastSpeech, neural vocoders, end-to-end and zero-shot voice models, evaluation, and voice-cloning ethics.Intermediate5 min
  29. 29Efficient Transformers: Sparse Attention, Linear Attention and FlashAttentionSelf-attention's quadratic cost limits context length. We survey sparse and local attention, low-rank and kernel-based linear attention, IO-aware FlashAttention, and alternatives such as state-space models — with their trade-offs.Advanced6 min