✨ Generative AI & LLMs · Lecture 14 of 30

Direct Preference Optimisation (DPO) and Beyond

DPO aligns language models to preferences with a simple classification-style loss — no reward model, no RL loop. We derive DPO from the KL-regularised RLHF objective, implement it, and survey variants and practical considerations.

RLHF works, but it is complex: train a reward model, then run unstable reinforcement learning with four large models in memory. In 2023, Rafailov, Sharma, Mitchell and colleagues showed that the same objective can be optimised directly on preference data with a simple loss. Direct Preference Optimisation (DPO) quickly became one of the most widely used alignment methods, especially for open models.

Starting point: the RLHF objective#

$$ \max_{\pi}\;\mathbb{E}_{x, y \sim \pi}\big[r(x, y)\big] - \beta\,D_{\text{KL}}\big(\pi(\cdot \mid x)\,\|\,\pi_{\text{ref}}(\cdot \mid x)\big) $$

This KL-regularised problem has a known closed-form optimal policy:

$$ \pi^*(y \mid x) = \frac{1}{Z(x)}\,\pi_{\text{ref}}(y \mid x)\exp\left(\frac{1}{\beta}r(x, y)\right) $$

where $Z(x)$ is a (intractable) normalising constant.

The key insight: rearrange for the reward#

Solve for $r$:

$$ r(x, y) = \beta\log\frac{\pi^*(y \mid x)}{\pi_{\text{ref}}(y \mid x)} + \beta\log Z(x) $$

Every reward function corresponds to a policy, and vice versa — "your language model is secretly a reward model". Now substitute this into the Bradley–Terry preference model. The troublesome $Z(x)$ appears in both rewards and cancels in their difference:

$$ P(y_w \succ y_l \mid x) = \sigma\left(\beta\log\frac{\pi(y_w \mid x)}{\pi_{\text{ref}}(y_w \mid x)} - \beta\log\frac{\pi(y_l \mid x)}{\pi_{\text{ref}}(y_l \mid x)}\right) $$

The DPO loss#

Maximise the likelihood of the observed preferences directly with respect to the policy parameters:

$$ \mathcal{L}_{\text{DPO}}(\theta) = -\mathbb{E}_{(x, y_w, y_l)}\left[\log\sigma\left(\beta\log\frac{\pi_\theta(y_w \mid x)}{\pi_{\text{ref}}(y_w \mid x)} - \beta\log\frac{\pi_\theta(y_l \mid x)}{\pi_{\text{ref}}(y_l \mid x)}\right)\right] $$

It looks like logistic regression: increase the (reference-relative) log-probability of the chosen response and decrease that of the rejected one. No sampling during training, no reward model, no value network — just forward passes of the policy and a frozen reference model on fixed preference pairs.

The gradient's intuition#

$$ \nabla_\theta\mathcal{L}_{\text{DPO}} = -\beta\,\mathbb{E}\Big[\sigma\big(\hat{r}_\theta(x, y_l) - \hat{r}_\theta(x, y_w)\big)\big(\nabla_\theta\log\pi_\theta(y_w \mid x) - \nabla_\theta\log\pi_\theta(y_l \mid x)\big)\Big] $$

with implicit rewards $\hat{r}_\theta = \beta\log\frac{\pi_\theta}{\pi_{\text{ref}}}$. Examples where the model currently ranks the pair wrongly receive larger weight — like a focusing mechanism.

Implementation#

python
import torch
import torch.nn.functional as F

def sequence_logprob(model, input_ids, labels):
    """Sum of log-probs of response tokens (labels == -100 elsewhere)."""
    logits = model(input_ids).logits[:, :-1]
    tgt = labels[:, 1:]
    logp = torch.gather(logits.log_softmax(-1), 2, tgt.clamp(min=0).unsqueeze(-1)).squeeze(-1)
    return (logp * (tgt != -100)).sum(-1)

def dpo_loss(pi_chosen, pi_rejected, ref_chosen, ref_rejected, beta=0.1):
    """Inputs: summed log-probs of chosen/rejected responses under policy and frozen reference."""
    logits = beta * ((pi_chosen - ref_chosen) - (pi_rejected - ref_rejected))
    loss = -F.logsigmoid(logits).mean()
    reward_acc = (logits > 0).float().mean()           # how often implicit reward ranks correctly
    return loss, reward_acc

print(dpo_loss(torch.tensor([-12.0, -30.0]), torch.tensor([-15.0, -28.0]),
               torch.tensor([-13.0, -29.0]), torch.tensor([-14.0, -29.0])))

In practice, libraries such as TRL provide DPOTrainer; data is a set of (prompt, chosen, rejected) triples, typically starting from an SFT model that is also used as the frozen reference.

Hyperparameters and practice#

  • $\beta$ (typically 0.01–0.5) controls deviation from the reference: smaller $\beta$ allows larger changes.
  • Low learning rates (e.g. $5\times10^{-7}$ to $5\times10^{-6}$ for full fine-tuning), one to a few epochs.
  • Preference data can be human labelled, AI labelled, or constructed by sampling several responses from the current model and ranking them (on-policy data tends to work better).
  • Monitor implicit reward accuracy, reward margins, and — crucially — generation quality on held-out prompts.

Known issues and variants#

  • Likelihood displacement: DPO can decrease the absolute probability of both chosen and rejected responses while increasing their gap, sometimes harming quality.
  • Overfitting to preference data and length exploitation (preferring longer responses) — regularisation or length-controlled variants help.
  • Offline: DPO learns from a fixed dataset; iterative/online DPO regenerates preference pairs with the current model.
MethodIdea
IPOAdds a regulariser that prevents overfitting when preferences are deterministic
KTOUses unpaired "good" / "bad" labels, inspired by prospect theory
ORPOCombines SFT and preference optimisation in one stage, no reference model
SimPOUses length-normalised log-probability as implicit reward, no reference model
Online / iterative DPOSamples new responses from the current policy each round

DPO vs RLHF#

PPO-based RLHFDPO
Reward modelExplicitImplicit
Sampling during trainingYes (on-policy)No (offline pairs)
Models in memoryPolicy, reference, reward, valuePolicy, reference
Stability and simplicityHarderEasier
Exploration beyond dataYesLimited (unless iterative)

Both remain in use; many modern pipelines combine SFT, preference optimisation (DPO-style) and RL with verifiable rewards.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

✨ Generative AI & LLMs

RLHF: Reinforcement Learning from Human Feedback

RLHF aligns language models with human preferences using a learned reward model and reinforcement learning. We derive the Bradley–Terry reward model, the KL-regularised objective optimised with PPO, and discuss reward hacking and limitations.

Advanced⏱ 5 min#203
✨ Generative AI & LLMs

Large Language Models: What They Are and How They Are Built

A map of large language models — the transformer backbone, the training pipeline from pretraining to alignment, what capabilities emerge, how they are served and used, and their fundamental limitations.

Beginner⏱ 5 min#199
✨ Generative AI & LLMs

Prompt Engineering: Getting Reliable Results from LLMs

Practical, evidence-based techniques for prompting LLMs — clear instructions, context, examples, output formats, decomposition, and systematic evaluation — plus prompt injection risks.

Beginner⏱ 5 min#205