"Roger has 5 tennis balls. He buys 2 more cans of 3 balls each. How many does he have?" Asked to answer immediately, early large models often got such problems wrong. Asked to reason step by step, they got many right. This observation launched a line of research that transformed how LLMs solve problems — culminating in "reasoning models" trained to think at length before answering.
Chain-of-thought prompting#
Wei et al. (2022) showed that providing few-shot examples with worked reasoning ("Roger started with 5 balls. 2 cans of 3 balls is 6. 5 + 6 = 11. The answer is 11.") greatly improved performance of large models on arithmetic, commonsense and symbolic reasoning benchmarks. The benefit appeared mainly in sufficiently large models.
Kojima et al. (2022) found that simply appending "Let's think step by step" — zero-shot chain-of-thought (CoT) — also produced large gains.
Why might it help?#
- More computation: each generated token is another forward pass; intermediate steps give the model more "serial computation" per problem.
- Decomposition: complex problems become a sequence of simpler predictions, each well supported by training data.
- Working memory: intermediate results are written into the context, where they can be attended to later.
Self-consistency#
Wang et al. (2023): sample many reasoning chains with temperature > 0, extract each final answer, and take the majority vote. Different reasoning paths that converge on the same answer are more likely correct. Self-consistency substantially improved accuracy on maths benchmarks over single greedy chains.
from collections import Counter
import re
def self_consistency(generate, question, n=10):
"""generate(prompt, temperature) -> text containing 'The answer is X.'"""
prompt = f"Q: {question}\nA: Let's think step by step."
answers = []
for _ in range(n):
text = generate(prompt, temperature=0.8)
m = re.search(r"answer is\s*([-\d.,]+)", text)
if m:
answers.append(m.group(1).rstrip(".").replace(",", ""))
if not answers:
return None, 0.0
best, votes = Counter(answers).most_common(1)[0]
return best, votes / len(answers) # answer and agreement (a rough confidence)The agreement rate also serves as a useful — if imperfect — confidence signal.
Beyond linear chains#
- Least-to-most prompting: first decompose a problem into sub-questions, then solve them in order.
- Tree of Thoughts (Yao et al., 2023): explore multiple partial reasoning paths as a tree, evaluate them, and backtrack — search over thoughts, helpful for puzzles and planning.
- Program-aided reasoning (PAL, Program of Thoughts): have the model write code for the computation and execute it — exact arithmetic instead of error-prone mental maths.
- ReAct (Yao et al., 2023): interleave reasoning with actions (search, tool calls) and observations — the basis of many agents.
- Self-refinement and verification: ask the model to critique and revise its answer, or check each step; effectiveness varies, and models are often poor at finding their own errors without external feedback.
Reasoning models and test-time compute#
A major development (from 2024) was training models with reinforcement learning to produce long chains of thought before answering, rewarded mainly by whether the final answer is correct (e.g. on maths problems with checkable answers and code with tests). Such reasoning models learned behaviours such as trying alternative approaches, checking intermediate results and backtracking. Several labs reported that accuracy on hard maths, science and coding benchmarks improves as the model is allowed to think longer — a new axis of scaling: test-time compute. Open research reports (e.g. DeepSeek-R1) described how RL with verifiable rewards alone could elicit such long reasoning.
Test-time compute can also be spent by sampling many candidates and selecting with a verifier or reward model (best-of-$N$), or by search guided by process reward models that score intermediate steps.
Is the chain of thought faithful?#
Practical guidance#
- For multi-step problems (maths, logic, planning, analysis), ask for step-by-step reasoning — or use a reasoning model.
- For simple lookups or classification, CoT adds cost and latency with little gain.
- Use tools (calculators, code execution, retrieval) for exact computation and facts.
- Use self-consistency or verification when correctness matters and cost allows.
- Keep reasoning separate from the final answer (e.g. a clear "Final answer:" line) so it can be parsed.