✨ Generative AI & LLMs · Lecture 16 of 30

Chain-of-Thought and Reasoning in Language Models

Asking models to show intermediate steps dramatically improves reasoning. We cover chain-of-thought prompting, self-consistency, tree search, program-aided reasoning, reasoning models trained with RL, test-time compute and the faithfulness question.

"Roger has 5 tennis balls. He buys 2 more cans of 3 balls each. How many does he have?" Asked to answer immediately, early large models often got such problems wrong. Asked to reason step by step, they got many right. This observation launched a line of research that transformed how LLMs solve problems — culminating in "reasoning models" trained to think at length before answering.

Chain-of-thought prompting#

Wei et al. (2022) showed that providing few-shot examples with worked reasoning ("Roger started with 5 balls. 2 cans of 3 balls is 6. 5 + 6 = 11. The answer is 11.") greatly improved performance of large models on arithmetic, commonsense and symbolic reasoning benchmarks. The benefit appeared mainly in sufficiently large models.

Kojima et al. (2022) found that simply appending "Let's think step by step" — zero-shot chain-of-thought (CoT) — also produced large gains.

Why might it help?#

  • More computation: each generated token is another forward pass; intermediate steps give the model more "serial computation" per problem.
  • Decomposition: complex problems become a sequence of simpler predictions, each well supported by training data.
  • Working memory: intermediate results are written into the context, where they can be attended to later.

Self-consistency#

Wang et al. (2023): sample many reasoning chains with temperature > 0, extract each final answer, and take the majority vote. Different reasoning paths that converge on the same answer are more likely correct. Self-consistency substantially improved accuracy on maths benchmarks over single greedy chains.

python
from collections import Counter
import re

def self_consistency(generate, question, n=10):
    """generate(prompt, temperature) -> text containing 'The answer is X.'"""
    prompt = f"Q: {question}\nA: Let's think step by step."
    answers = []
    for _ in range(n):
        text = generate(prompt, temperature=0.8)
        m = re.search(r"answer is\s*([-\d.,]+)", text)
        if m:
            answers.append(m.group(1).rstrip(".").replace(",", ""))
    if not answers:
        return None, 0.0
    best, votes = Counter(answers).most_common(1)[0]
    return best, votes / len(answers)             # answer and agreement (a rough confidence)

The agreement rate also serves as a useful — if imperfect — confidence signal.

Beyond linear chains#

  • Least-to-most prompting: first decompose a problem into sub-questions, then solve them in order.
  • Tree of Thoughts (Yao et al., 2023): explore multiple partial reasoning paths as a tree, evaluate them, and backtrack — search over thoughts, helpful for puzzles and planning.
  • Program-aided reasoning (PAL, Program of Thoughts): have the model write code for the computation and execute it — exact arithmetic instead of error-prone mental maths.
  • ReAct (Yao et al., 2023): interleave reasoning with actions (search, tool calls) and observations — the basis of many agents.
  • Self-refinement and verification: ask the model to critique and revise its answer, or check each step; effectiveness varies, and models are often poor at finding their own errors without external feedback.

Reasoning models and test-time compute#

A major development (from 2024) was training models with reinforcement learning to produce long chains of thought before answering, rewarded mainly by whether the final answer is correct (e.g. on maths problems with checkable answers and code with tests). Such reasoning models learned behaviours such as trying alternative approaches, checking intermediate results and backtracking. Several labs reported that accuracy on hard maths, science and coding benchmarks improves as the model is allowed to think longer — a new axis of scaling: test-time compute. Open research reports (e.g. DeepSeek-R1) described how RL with verifiable rewards alone could elicit such long reasoning.

Test-time compute can also be spent by sampling many candidates and selecting with a verifier or reward model (best-of-$N$), or by search guided by process reward models that score intermediate steps.

Is the chain of thought faithful?#

Practical guidance#

  • For multi-step problems (maths, logic, planning, analysis), ask for step-by-step reasoning — or use a reasoning model.
  • For simple lookups or classification, CoT adds cost and latency with little gain.
  • Use tools (calculators, code execution, retrieval) for exact computation and facts.
  • Use self-consistency or verification when correctness matters and cost allows.
  • Keep reasoning separate from the final answer (e.g. a clear "Final answer:" line) so it can be parsed.
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

✨ Generative AI & LLMs

Prompt Engineering: Getting Reliable Results from LLMs

Practical, evidence-based techniques for prompting LLMs — clear instructions, context, examples, output formats, decomposition, and systematic evaluation — plus prompt injection risks.

Beginner⏱ 5 min#205
✨ Generative AI & LLMs

In-Context Learning: How LLMs Learn from Prompts

Large language models can perform new tasks from a few examples in the prompt without weight updates. We examine what in-context learning is, what influences it, theories of how it works, and its practical limits.

Intermediate⏱ 5 min#207
✨ Generative AI & LLMs

Direct Preference Optimisation (DPO) and Beyond

DPO aligns language models to preferences with a simple classification-style loss — no reward model, no RL loop. We derive DPO from the KL-regularised RLHF objective, implement it, and survey variants and practical considerations.

Advanced⏱ 5 min#204