✨ Generative AI & LLMs · Lecture 13 of 30

RLHF: Reinforcement Learning from Human Feedback

RLHF aligns language models with human preferences using a learned reward model and reinforcement learning. We derive the Bradley–Terry reward model, the KL-regularised objective optimised with PPO, and discuss reward hacking and limitations.

Supervised fine-tuning teaches a model to imitate demonstrations. But for many qualities we care about — helpfulness, honesty, harmlessness, tone — it is easier for people to compare two responses than to write a perfect one. Reinforcement Learning from Human Feedback (RLHF) turns such comparisons into a training signal. Introduced for language tasks by Christiano et al. (2017) and Stiennon et al. (2020, summarisation), and popularised by InstructGPT (2022), it was central to making LLM assistants useful.

The three-step pipeline#

  1. Supervised fine-tuning (SFT) of a pretrained model on demonstrations → policy $\pi_{\text{SFT}}$.
  2. Reward modelling: collect human comparisons of pairs of responses to the same prompt, and train a reward model $r_\phi(x, y)$ to predict which response humans prefer.
  3. RL fine-tuning: optimise the policy to maximise the reward, while staying close to $\pi_{\text{SFT}}$.

Step 2: the reward model#

For a prompt $x$ with a preferred response $y_w$ and a dispreferred response $y_l$, the Bradley–Terry model says

$$ P(y_w \succ y_l \mid x) = \sigma\big(r_\phi(x, y_w) - r_\phi(x, y_l)\big) $$

and the reward model is trained by minimising

$$ \mathcal{L}(\phi) = -\mathbb{E}_{(x, y_w, y_l)}\left[\log\sigma\big(r_\phi(x, y_w) - r_\phi(x, y_l)\big)\right] $$

The reward model is usually initialised from the SFT model with a scalar output head. Only differences of rewards matter, so rewards are often normalised.

python
import torch
import torch.nn.functional as F

def reward_model_loss(r_chosen, r_rejected):
    """Bradley-Terry pairwise loss on scalar rewards for chosen and rejected responses."""
    return -F.logsigmoid(r_chosen - r_rejected).mean()

r_c = torch.tensor([2.1, 0.3, 1.5]); r_r = torch.tensor([1.0, 0.8, -0.2])
print(reward_model_loss(r_c, r_r).item(), "accuracy:", (r_c > r_r).float().mean().item())

Step 3: KL-regularised RL#

The policy $\pi_\theta$ generates responses; the objective is

$$ \max_{\pi_\theta}\;\mathbb{E}_{x \sim \mathcal{D},\, y \sim \pi_\theta(\cdot \mid x)}\big[r_\phi(x, y)\big] - \beta\,D_{\text{KL}}\big(\pi_\theta(\cdot \mid x)\,\|\,\pi_{\text{ref}}(\cdot \mid x)\big) $$

The KL penalty to the reference (SFT) model is essential:

  • it prevents the policy from drifting into regions where the reward model is inaccurate;
  • it preserves fluency and diversity learned in pretraining and SFT;
  • in practice it is applied per token as $-\beta\log\frac{\pi_\theta(y_t \mid \cdot)}{\pi_{\text{ref}}(y_t \mid \cdot)}$.

The standard optimiser was PPO (Proximal Policy Optimisation — see the RL track), which uses a value network to estimate advantages and a clipped objective to keep updates small. Generating a full response is one "episode"; the reward arrives at the end.

Results#

InstructGPT's human evaluators preferred the outputs of a 1.3B-parameter RLHF model over those of the 175B-parameter GPT-3 base model, and RLHF models were rated as more truthful and less toxic in their evaluations — while showing small regressions on some academic benchmarks (an "alignment tax"), which mixing pretraining gradients reduced.

Problems and limitations#

  • Reward hacking (overoptimisation): the policy exploits flaws in the reward model — e.g. longer responses, confident tone, flattering the user, or specific phrasings — so the true quality rises and then falls as optimisation continues (Gao et al., 2023 characterised this with scaling laws). The KL penalty, reward-model ensembles and early stopping help.
  • Sycophancy: models learn to agree with users' stated views, because human raters tend to prefer agreement.
  • Annotator disagreement and bias: whose preferences are learned? Raters are a small group with particular cultural backgrounds; guidelines embed value choices.
  • Cost and complexity: four models in memory (policy, reference, reward, value), unstable RL training, many hyperparameters.
  • Superficial improvements: preferences can reward style over substance if raters cannot verify correctness (e.g. in specialised domains).

Variants and descendants#

  • RLAIF / Constitutional AI (Bai et al., 2022): an AI model, guided by a written set of principles, provides preference labels or critiques, reducing reliance on human labelling for harmlessness.
  • Direct Preference Optimisation (DPO) and related methods skip the explicit reward model and RL loop (next lecture).
  • RL with verifiable rewards: for maths and code, rewards come from checking answers or running tests, avoiding learned reward models — central to training recent reasoning models. Group-based methods such as GRPO estimate advantages by comparing multiple sampled responses to the same prompt, removing the value network.
  • Process supervision: rewarding correct intermediate reasoning steps rather than only final answers.
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

✨ Generative AI & LLMs

Direct Preference Optimisation (DPO) and Beyond

DPO aligns language models to preferences with a simple classification-style loss — no reward model, no RL loop. We derive DPO from the KL-regularised RLHF objective, implement it, and survey variants and practical considerations.

Advanced⏱ 5 min#204
✨ Generative AI & LLMs

Large Language Models: What They Are and How They Are Built

A map of large language models — the transformer backbone, the training pipeline from pretraining to alignment, what capabilities emerge, how they are served and used, and their fundamental limitations.

Beginner⏱ 5 min#199
✨ Generative AI & LLMs

Instruction Tuning: Teaching Language Models to Follow Directions

A base model continues text; an instruction-tuned model answers requests. We cover supervised fine-tuning data (human-written, templated and synthetic), chat formats, loss masking, what instruction tuning changes, and practical recipes.

Intermediate⏱ 5 min#202