✨ Generative AI & LLMs · Lecture 9 of 30

Large Language Models: What They Are and How They Are Built

A map of large language models — the transformer backbone, the training pipeline from pretraining to alignment, what capabilities emerge, how they are served and used, and their fundamental limitations.

Large Language Models (LLMs) are the most visible AI technology of our time. They draft emails, explain physics, write code, translate, summarise reports and hold conversations in dozens of languages. This lecture gives you a clear map of how they are built and what they can and cannot do; subsequent lectures examine each stage in depth.

What is an LLM?#

An LLM is (almost always) a decoder-only transformer trained to predict the next token, with billions to hundreds of billions of parameters, trained on trillions of tokens of text (and often code, images and other modalities). Given a sequence of tokens, it outputs a probability distribution over the next token; repeated sampling generates text.

The training pipeline#

StageDataObjectiveResult
1. PretrainingTrillions of tokens: web, books, code, papers, multilingual textNext-token predictionBase model: broad knowledge and skills, but just continues text
2. Supervised fine-tuning (SFT)Thousands to millions of instruction–response demonstrationsNext-token prediction on responsesFollows instructions, adopts assistant format
3. Preference optimisationHuman (or AI) comparisons of responsesRLHF, DPO or similarMore helpful, honest, harmless; better style
4. Reasoning / RL training (in recent models)Problems with verifiable answers (maths, code)Reinforcement learning on outcomesLonger, more reliable step-by-step reasoning
5. Safety evaluation and red-teamingAdversarial prompts, expert testing—Identified risks, mitigations, usage policies

Pretraining dominates compute (often thousands of GPUs for weeks or months); post-training stages are far cheaper but shape behaviour dramatically.

Capabilities#

LLMs can:

  • follow instructions and hold multi-turn conversations;
  • perform in-context learning from examples in the prompt;
  • write and debug code;
  • summarise, translate, classify and extract information;
  • reason step by step on many problems, especially when prompted or trained to do so;
  • use tools — search, calculators, code interpreters, APIs — when given function-calling interfaces;
  • process images, audio and documents (multimodal models).

Limitations#

  • Hallucination: fluent, confident statements that are false or unsupported.
  • Knowledge cut-off: knowledge is frozen at training time unless connected to retrieval or tools.
  • Reasoning brittleness: errors on problems requiring long exact computation, careful counting, or novel structures; sensitivity to phrasing.
  • Context limits: finite context windows and imperfect use of long contexts.
  • Bias and representational harms absorbed from training data.
  • Security issues: prompt injection, jailbreaks, data leakage.
  • Cost and latency of large models; environmental footprint.

How LLMs are used#

  1. Prompting — zero-shot or few-shot instructions (prompt engineering).
  2. Retrieval-augmented generation (RAG) — ground answers in your documents.
  3. Fine-tuning — adapt to a domain, format or task, often with parameter-efficient methods (LoRA).
  4. Agents — LLMs that plan and call tools in loops to complete tasks.
  5. Embedding models — related encoders for search and clustering.

Open and closed models#

  • Closed/proprietary models are accessed via APIs; typically strongest capabilities; data handling governed by provider terms.
  • Open-weight models can be downloaded and run locally or on private servers; important for privacy, customisation, cost control and research. They vary widely in licence terms.

For organisations handling sensitive personal data — health records, refugee case files — the ability to run models on controlled infrastructure can be decisive.

A first API-style example#

python
# Using an open model locally with Hugging Face transformers
from transformers import pipeline

chat = pipeline("text-generation", model="Qwen/Qwen2.5-0.5B-Instruct")
messages = [
    {"role": "system", "content": "You are a concise tutor for university students."},
    {"role": "user", "content": "Explain overfitting in two sentences with an everyday analogy."},
]
out = chat(messages, max_new_tokens=120, do_sample=False)
print(out[0]["generated_text"][-1]["content"])

The chat template converts role-tagged messages into the special token format the model was fine-tuned on — using the wrong template degrades performance.

Thinking clearly about LLMs#

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

✨ Generative AI & LLMs

Pretraining LLMs: Data Pipelines, Objectives and Infrastructure

What actually goes into pretraining an LLM? We cover data sourcing, filtering, deduplication, mixture design, tokenisation, the training objective, stability tricks, infrastructure and evaluation during pretraining.

Advanced⏱ 5 min#201
✨ Generative AI & LLMs

RLHF: Reinforcement Learning from Human Feedback

RLHF aligns language models with human preferences using a learned reward model and reinforcement learning. We derive the Bradley–Terry reward model, the KL-regularised objective optimised with PPO, and discuss reward hacking and limitations.

Advanced⏱ 5 min#203
✨ Generative AI & LLMs

Direct Preference Optimisation (DPO) and Beyond

DPO aligns language models to preferences with a simple classification-style loss — no reward model, no RL loop. We derive DPO from the KL-regularised RLHF objective, implement it, and survey variants and practical considerations.

Advanced⏱ 5 min#204