SFT teaches the model to imitate demonstrations. Alignment methods teach it to prefer better answers over worse ones—helpfulness, harmlessness, and style trade-offs that are hard to encode as a single gold string. This lesson is a map of RLHF and DPO for engineers who need to know when to reach for them.
Humans often cannot write the perfect answer, but they can say "A is better than B." Preference learning turns those comparisons into training signal.
SFT: learn p(y | x) from chosen y*
Prefs: learn to rank y_w > y_l given x
RLHF (Reinforcement Learning from Human Feedback) trains a reward model on preferences, then optimizes the policy (the LLM) to score high under that reward while staying close to an SFT reference.
DPO (Direct Preference Optimization) skips the explicit RL loop and reward model: it updates the policy directly from preference pairs with a closed-form objective relative to a reference model.
x, winner y_w and loser y_l (human or strong-AI labeled).High level:
r(x, y) predicts human preference.maximize E[r(x, y)] - beta * KL(pi_theta || pi_ref)
Engineering cost: sampling trajectories, unstable training, more moving parts (RM + policy + value/critic depending on stack).
DPO reparameterizes the problem so that increasing the likelihood gap between y_w and y_l (relative to the reference) improves the preference objective—no separate RM training loop in the common recipe.
increase log pi(y_w|x) - log pi(y_l|x)
relative to the same gap under pi_ref
(scaled by beta)
Teams like DPO when they want preference gains with a simpler trainer. Quality still hinges on preference data.
| Situation | Lean toward |
|---|---|
| Need format/style imitation only | SFT |
| Clear pairwise judgments at scale | DPO or RLHF |
| Complex multi-objective reward already modeled | RLHF-style stacks |
| Small team, limited RL ops experience | DPO (or SFT+better data) first |
| Safety-critical nuanced refusals | Preference data + strong eval; not vibes |
Toy preference update: push up winner score, push down loser, keep a reference gap in mind.
import math
def sigmoid(z: float) -> float:
return 1.0 / (1.0 + math.exp(-z))
# Log-probs under policy and frozen reference (toy scalars)
logp = {"win": -2.0, "lose": -2.1}
logp_ref = {"win": -2.0, "lose": -2.0}
beta = 0.1
lr = 0.5
for step in range(5):
# DPO-like advantage of win vs lose vs reference
delta = (logp["win"] - logp["lose"]) - (logp_ref["win"] - logp_ref["lose"])
# Gradient signal: increase delta
loss = -math.log(sigmoid(beta * delta) + 1e-9)
# Finite-difference style updates on logp
logp["win"] += lr * beta * (1 - sigmoid(beta * delta))
logp["lose"] -= lr * beta * (1 - sigmoid(beta * delta))
print(f"step {step} delta={delta:.3f} loss={loss:.3f} logp={logp}")
Conceptual trainer wiring:
# pairs: list of {prompt, chosen, rejected}
# for batch in pairs:
# loss = dpo_loss(policy, ref_policy, batch, beta=0.1)
# loss.backward(); optimizer.step()
# Never required for this course: running PPO on a GPU cluster.
Good pairs share the same prompt and differ on the dimension you care about (correctness, tone, refusal). Bad pairs mix unrelated axes (one answer wrong, the other just longer). Write rater guidelines with examples of ties, and measure inter-rater agreement. If agreement is poor, fix the guide before training.
Ship SFT when the main bugs are format and imitation. Reach for preferences when raters systematically prefer B over A even though both are "valid," or when safety policy needs graded judgment. Many teams do SFT → DPO → evaluate; RLHF stacks appear when they already invest in reward modeling infrastructure.
The reference model is a leash. Without it, optimizers invent high-reward nonsense: endless apologies, empty hedging, or keyword stuffing that fooled the reward model. If outputs get strangely verbose or sycophantic after preference tuning, strengthen the leash (higher beta / KL) and revisit the preference labels.
RLHF learns a reward model and optimizes the policy with RL; DPO learns from preference pairs more directly—both refine an SFT model toward ranked human (or proxy) judgments.