Cookbook · DPO
DPO (Direct Preference Optimization)
3 mindporlhfpreference-optimizationfine-tuning
The idea, in one analogy
Imagine training a dog with a scale you never built: instead of measuring exactly how good a trick was and rewarding proportionally, you just show the dog two attempts side by side and say "that one, not that one." DPO trains a language model the same way: instead of building a separate reward model to score responses and then running reinforcement learning against that score (see RLHF), it trains directly on pairs of responses, one preferred over the other, adjusting the model to make the preferred one relatively more likely.
Why skip the reward model and RL loop
RLHF's standard recipe is two stages: train a reward model on human preference comparisons, then run reinforcement learning (usually PPO) to optimize the policy against that reward, with a KL penalty to keep it from drifting too far from the original model. Both stages are individually finicky: reward models can be exploited by the policy finding outputs that score well without being genuinely better ("reward hacking"), and RL training is notoriously sensitive to hyperparameters and prone to instability.
DPO's insight is mathematical: for a specific, common choice of reward model (one where the reward is derived directly from the policy's own probabilities, relative to a reference model), the RLHF objective has a closed-form solution. That means you can skip training a separate reward model and skip the RL loop entirely, and instead optimize a single, simple classification-style loss directly on the preference pairs, one that provably has the same optimum RLHF's more complicated two-stage process was aiming for.
What you need to run it
A dataset of (prompt, winning response, losing response) triples, and two copies of the model in memory during training: the policy being updated and the frozen reference. That reference copy is the main resource cost DPO adds beyond ordinary fine-tuning; there's no reward model to train or serve separately, and no RL rollout loop generating fresh responses during training the way PPO-based RLHF does.
Where to look further
- Rafailov et al., "Direct Preference Optimization: Your Language Model is Secretly a Reward Model": the original paper, including the derivation connecting the RLHF objective to this closed-form loss.
- Hugging Face's TRL library: includes a maintained
DPOTrainerimplementation. - RLHF and GRPO: the two main alternatives, each with different tradeoffs around reward models and RL machinery.