Personalize this lesson
Adapt explanations and teaching visuals to your background and preferred voice.
The last chapter left you with a resumable supervised fine-tuning (SFT) job. policy-sft-v4 can continue, initialize, or export, and you already know which adaptation method to launch. Once that checkpoint answers, a new training problem shows up: two fluent replies to the same access-policy prompt. One cites policy P-7 and opens a review ticket. The other grants admin access for the night. Which one should score higher?
RLHF (Reinforcement Learning from Human Feedback) diagrams often squash that question into a box labeled reward model, then jump to Proximal Policy Optimization (PPO). The box is its own training project. If it learns the wrong shortcuts, policy optimization will amplify them.[1]
A scalar judge, not a generator
A reward model doesn't write the next token. It scores text that already exists.
Given a prompt x and two candidate answers:
y+chosen by the labeler (open a review ticket, cite P-7)y-rejected by the labeler (grant admin tonight)
the model should assign:
1r(x, y+) > r(x, y-)That scalar is useful in two different ways later:
- as a reward signal for PPO-style RLHF[1]
- as an inspectable ranking score when you want to compare many policy outputs, including best-of-N sampling
Explicit reward models still earn their keep even though Direct Preference Optimization (DPO) can skip them for offline preference optimization.[2] Train that scalar first, then distrust it, before you plug it into either use.
Preference pairs that can actually train
The core supervision format is a preference pair: same prompt, two different completions, one definite ranking.
Standard format
1{
2 "prompt": "User requests temporary admin access for a migration. What should the assistant do?",
3 "chosen": "Open an access-review ticket and cite policy P-7 before approval.",
4 "rejected": "Grant admin for tonight and ask the user to clean it up tomorrow."
5}Conversational format
1{
2 "chosen": [
3 {"role": "user", "content": "User requests temporary admin access for a migration. What should the assistant do?"},
4 {"role": "assistant", "content": "Open an access-review ticket and cite policy P-7 before approval."}
5 ],
6 "rejected": [
7 {"role": "user", "content": "User requests temporary admin access for a migration. What should the assistant do?"},
8 {"role": "assistant", "content": "Grant admin for tonight and ask the user to clean it up tomorrow."}
9 ]
10}Hugging Face TRL's RewardTrainer accepts both standard and conversational preference formats and can apply the model's chat template during tokenization.[3]
A pair contract before training
A row isn't ready just because it has chosen and rejected columns. For every binary pair, enforce:
- the same prompt, system message, tool context, and rendering template on both candidates
- two different candidate answers and a definite preference label
- separate handling for ties, abstentions, and low-agreement labels
- provenance such as source prompt, generator checkpoint, sampling settings, and labeling batch
That last field matters for splitting. If one prompt generated several candidates, its comparisons are near-duplicates. Put the entire prompt group in train or in evaluation, never both.
The audit below keeps only row a. Identical candidates, mismatched prompts, and ties don't enter ordinary pairwise ranking training.
1from collections import Counter
2
3pairs = [
4 {"id": "a", "prompt_left": "admin?", "prompt_right": "admin?", "chosen": "Escalate.", "rejected": "Approve.", "label": "chosen"},
5 {"id": "b", "prompt_left": "admin?", "prompt_right": "admin?", "chosen": "Escalate.", "rejected": "Escalate.", "label": "chosen"},
6 {"id": "c", "prompt_left": "admin?", "prompt_right": "source?", "chosen": "Escalate.", "rejected": "Cite source.", "label": "chosen"},
7 {"id": "d", "prompt_left": "source?", "prompt_right": "source?", "chosen": "Cite.", "rejected": "Refuse.", "label": "tie"},
8]
9
10def rejection_reason(pair):
11 if pair["label"] != "chosen":
12 return "tie_or_abstention"
13 if pair["prompt_left"] != pair["prompt_right"]:
14 return "context_mismatch"
15 if pair["chosen"] == pair["rejected"]:
16 return "identical_candidates"
17 return None
18
19reasons = Counter(reason for pair in pairs if (reason := rejection_reason(pair)))
20kept = [pair["id"] for pair in pairs if rejection_reason(pair) is None]
21
22print(f"kept={kept}")
23print(f"rejected={dict(sorted(reasons.items()))}")1kept=['a']
2rejected={'context_mismatch': 1, 'identical_candidates': 1, 'tie_or_abstention': 1}A grouped split is the next check. Rows a-b and a-c share prompt support-17, so they move together.
1pairs = [
2 {"pair_id": "a-b", "prompt_id": "support-17"},
3 {"pair_id": "a-c", "prompt_id": "support-17"},
4 {"pair_id": "d-e", "prompt_id": "safety-04"},
5 {"pair_id": "f-g", "prompt_id": "access-09"},
6]
7eval_prompt_ids = {"support-17"}
8
9train = [pair for pair in pairs if pair["prompt_id"] not in eval_prompt_ids]
10evaluation = [pair for pair in pairs if pair["prompt_id"] in eval_prompt_ids]
11overlap = {pair["prompt_id"] for pair in train} & {pair["prompt_id"] for pair in evaluation}
12
13assert not overlap
14print(f"train_pairs={len(train)} eval_pairs={len(evaluation)}")
15print(f"prompt_overlap={sorted(overlap)}")1train_pairs=2 eval_pairs=2
2prompt_overlap=[]Prompt groups also change how you train, not only how you split. InstructGPT asked labelers to rank to completions per prompt, which produces pairwise comparisons. If you shuffle those pairs into the dataset independently, each completion can contribute gradient updates in one epoch, and the reward model overfits. Their fix was to treat every comparison from one prompt as a single batch element and average the loss over the pairs.[1]
Last token, one number
Most practical reward models aren't built from scratch. You start from a pretrained or SFT checkpoint, drop the token-unembedding head, and attach a one-unit score head. InstructGPT did exactly that: replace the unembedding layer with a projection to a scalar.[1] Conceptually:
1tokens | transformer hidden states | sequence representation | one scalar rewardFor a decoder-only LM, that representation is often the final non-padding position. Reward modeling then feels like sequence classification with pairwise labels rather than generation. The output is one number per candidate response, not the next-token distribution.
The rendered sequence is part of the contract. Keep chat-template and end-of-sequence conventions consistent with the policy you'll score. Don't silently train on answers whose decisive ending was truncated: current TRL RewardConfig.max_length (default 1024) drops a pair when either candidate exceeds that limit after tokenization.[3]
1max_length = 1024
2pairs = [
3 {"id": "fits", "chosen_tokens": 412, "rejected_tokens": 390},
4 {"id": "chosen_too_long", "chosen_tokens": 1088, "rejected_tokens": 401},
5 {"id": "rejected_too_long", "chosen_tokens": 288, "rejected_tokens": 1030},
6]
7
8kept = [p["id"] for p in pairs if max(p["chosen_tokens"], p["rejected_tokens"]) <= max_length]
9dropped = [p["id"] for p in pairs if p["id"] not in kept]
10
11print(f"kept={kept}")
12print(f"dropped_instead_of_truncated={dropped}")1kept=['fits']
2dropped_instead_of_truncated=['chosen_too_long', 'rejected_too_long']A pair that scores is only useful if the loss on those two numbers matches the label. That's the next question.
Bradley-Terry from a worked margin
Start with the access-policy pair. Suppose the ticket-and-cite answer gets reward 1.8 and the unsupported approval gets 0.7. The margin is 1.8 - 0.7 = 1.1. A positive margin means the reward model prefers the chosen answer. The Bradley-Terry model turns that margin into a preference probability: .

where is the sigmoid. InstructGPT trains the negative log of that probability, averaged over the comparisons in a prompt group:[1]
where is the chosen (winning) and rejected (losing) pair among the ranked completions for prompt . For a single binary pair, and the factor is 1, so the loss is just . If the chosen reward is much higher than the rejected reward, loss shrinks toward 0. If the model ranks them backwards, loss gets large.
The numerically stable form of is . For negative , rewrite it as so exp doesn't overflow.
Why does a larger positive reward margin make the Bradley-Terry loss smaller?
Answer
Because the loss is -log sigma(r(chosen) - r(rejected)). As the chosen-minus-rejected margin grows, the sigmoid term moves closer to 1, so the negative log shrinks toward 0.
The batch below uses the same three margins you'll see in the quiz: 1.7, -0.2, and 0.7. Pair accuracy is 2/3 because one pair is ranked backwards.
1from math import exp, log1p
2from statistics import mean
3
4chosen_rewards = [2.1, 0.8, 1.9]
5rejected_rewards = [0.4, 1.0, 1.2]
6margins = [a - b for a, b in zip(chosen_rewards, rejected_rewards)]
7
8def neg_log_sigmoid(z: float) -> float:
9 if z >= 0:
10 return log1p(exp(-z))
11 return -z + log1p(exp(z))
12
13loss = mean(neg_log_sigmoid(m) for m in margins)
14accuracy = mean(m > 0 for m in margins)
15
16print("margins:", [round(m, 3) for m in margins])
17print("reward_loss:", round(loss, 4))
18print("pair_accuracy:", round(accuracy, 4))1margins: [1.7, -0.2, 0.7]
2reward_loss: 0.4564
3pair_accuracy: 0.6667Sweeping the margin makes the same loss visible as a curve. Zero margin is a coin flip (log 2 ≈ 0.6931). A margin of -2 is a confident wrong ranking.
1from math import exp, log1p
2
3def neg_log_sigmoid(z: float) -> float:
4 if z >= 0:
5 return log1p(exp(-z))
6 return -z + log1p(exp(z))
7
8for margin in [-2.0, 0.0, 2.0]:
9 print(f"margin={margin:+.1f} loss={neg_log_sigmoid(margin):.4f}")1margin=-2.0 loss=2.1269
2margin=+0.0 loss=0.6931
3margin=+2.0 loss=0.1269That's the ranking core. Everything else in reward modeling is about whether the data and evaluation around that loss are strong enough to trust.
Multi-way rankings and a preference-strength margin
Binary chosen/rejected pairs are the default training unit. Production sets often produce completions per prompt. Two standard expansions:
- Pair expansion. InstructGPT's method: emit every ordered pair the ranking can score, then pack those pairs by prompt so one completion doesn't dominate the epoch.[1]
- Plackett-Luce. Model a full ranking as a product of successive choices rather than independent pairs. Use it when you have true multi-way rankings and want one joint likelihood. Pair expansion remains the simpler baseline when only sparse pairwise labels exist.
Llama 2 added a discrete margin inside the sigmoid when raters marked how much better the winner was (significantly better, slightly better, and so on):[4]
Positive requires the chosen score to beat the rejected score by more than before the loss saturates. Llama 3 later dropped that term after seeing diminishing returns once preference data scaled.[5] Treat as a hyperparameter you validate on held-out pair accuracy and fresh-policy human agreement, not as a free constant.
A temperature that multiplies the reward difference, , is multiplicative score scaling: it changes loss confidence and downstream RL strength without changing the additive invariance of pure ranking. Log any or next to the reward checkpoint.
What pair accuracy hides
TRL logs more than loss for a reason.[3]
| Metric | What it tells you | What it doesn't tell you |
|---|---|---|
| Pair accuracy | How often chosen beats rejected | Whether the ranking reason is policy-relevant |
| Mean margin | Typical r(chosen) - r(rejected) | Whether scores are calibrated for PPO |
| Mean / min / max reward | Drift or exploding scale | Whether humans would agree on new outputs |
| Gradient norm | Unstable updates | Shortcut features such as length |
| Held-out preference quality | Ranking on a static split | Ranking on the current policy's fresh outputs |
Loss can fall while the model memorizes easy stylistic cues that overfit the train pairs. InstructGPT already saw this: extra epochs quickly hurt validation loss even when the architecture was stable.[1]
Centering and calibration
Reward models are underdetermined up to an additive constant: adding the same number to every score leaves every margin and the Bradley-Terry loss unchanged. Scaling scores is different. It preserves a ranking but changes loss confidence and the strength of a reward signal consumed by an optimizer.
That matters operationally because:
- absolute reward level and score scale can drift over training
- PPO-style optimization is sensitive to reward scale
- long verbose answers can look better than they are if the model learned a shallow heuristic
InstructGPT fixed the additive ambiguity with a bias so labeler demonstrations scored mean zero before RL.[1] TRL exposes center_rewards_coefficient (recommended 0.01) as an auxiliary term that pushes batch rewards toward zero during training. It's a centering aid, not proof that reward magnitude is calibrated for policy optimization.[3]
1from math import exp, log1p
2from statistics import mean
3
4chosen = [1.2, 0.7]
5rejected = [0.2, 0.4]
6
7def pair_loss(left, right):
8 return mean(log1p(exp(-(a - b))) for a, b in zip(left, right))
9
10shifted = ([score + 10 for score in chosen], [score + 10 for score in rejected])
11scaled = ([score * 3 for score in chosen], [score * 3 for score in rejected])
12
13print(f"base_loss={pair_loss(chosen, rejected):.4f}")
14print(f"shifted_loss={pair_loss(*shifted):.4f}")
15print(f"scaled_loss={pair_loss(*scaled):.4f}")1base_loss=0.4338
2shifted_loss=0.4338
3scaled_loss=0.1949Shift leaves the loss alone. Scale doesn't. If you later feed these scores into PPO, the * 3 version is a stronger (and easier to overoptimize) signal even though the ranking is the same.
Audits that pair accuracy misses
Pair accuracy can look healthy while the reward model still learns a bad shortcut. This audit ranks all three preference pairs correctly, but every chosen answer is longer than its rejected pair.
1from statistics import mean
2
3pairs = [
4 {"chosen_reward": 1.8, "rejected_reward": 0.7, "chosen_tokens": 36, "rejected_tokens": 19},
5 {"chosen_reward": 2.4, "rejected_reward": 1.1, "chosen_tokens": 58, "rejected_tokens": 22},
6 {"chosen_reward": 1.6, "rejected_reward": 0.3, "chosen_tokens": 33, "rejected_tokens": 14},
7]
8
9margins = [row["chosen_reward"] - row["rejected_reward"] for row in pairs]
10accuracy = mean(margin > 0 for margin in margins)
11length_gaps = [row["chosen_tokens"] - row["rejected_tokens"] for row in pairs]
12
13print(f"pair_accuracy={accuracy:.2f}")
14print(f"mean_margin={mean(margins):.2f}")
15print(f"chosen_answers_longer={all(gap > 0 for gap in length_gaps)}")
16print("next_check=build length-matched eval pairs")1pair_accuracy=1.00
2mean_margin=1.23
3chosen_answers_longer=True
4next_check=build length-matched eval pairsAnnotator disagreement is another failure signal. A pair can be formatted correctly and still be weak supervision if raters don't agree about which completion is better. The toy gate below routes any disputed label to review. A real pipeline may adjudicate, weight, or keep disagreements on a dedicated evaluation slice.
1votes = {
2 "clear_safety": ["chosen", "chosen", "chosen"],
3 "style_only": ["chosen", "rejected", "chosen"],
4 "ambiguous_refusal": ["chosen", "rejected", "tie"],
5}
6
7for pair_id, labels in votes.items():
8 chosen_share = labels.count("chosen") / len(labels)
9 status = "train" if chosen_share == 1.0 else "review_or_hold_out"
10 print(f"{pair_id}: chosen_share={chosen_share:.2f} status={status}")1clear_safety: chosen_share=1.00 status=train
2style_only: chosen_share=0.67 status=review_or_hold_out
3ambiguous_refusal: chosen_share=0.33 status=review_or_hold_outLength-matched slices and rater agreement still evaluate the original pair distribution. The harder question is what happens when the policy starts writing answers that weren't in that distribution.
Fresh outputs are the real test
A reward model can fit the training pairs and still rank the current policy's fresh answers the wrong way. That second ranking is the one that matters before you optimize against the score.
As the policy improves, it starts producing answers unlike the ones in the original preference dataset. The reward model may then score confidently for the wrong reasons. Gao, Schulman, and Hilton measured this as overoptimization: a proxy reward you keep optimizing eventually stops tracking a gold preference signal, which is Goodhart's law with a learning curve.[6] In their synthetic setup the gold score rose, then fell, while the proxy kept climbing. A KL penalty increased proxy reward at a given KL, but it didn't improve the gold-reward frontier.
Watch for that gap in production:
- the learned reward rises
- held-out human preference on fresh policy outputs stops rising
- raters see longer, repetitive, or otherwise worse answers

If you don't monitor that gap, policy optimization can amplify the shortcut. The threshold below is illustrative; set release gates from your evaluation design and risk tolerance.
1evaluation = {
2 "static_held_out_pairs": {"accuracy": 0.92, "human_reviewed": False},
3 "fresh_policy_pairs": {"accuracy": 0.64, "human_reviewed": True},
4}
5minimum_fresh_accuracy = 0.80
6ppo_ready = evaluation["fresh_policy_pairs"]["accuracy"] >= minimum_fresh_accuracy
7
8print(f"static_accuracy={evaluation['static_held_out_pairs']['accuracy']:.2f}")
9print(f"fresh_accuracy={evaluation['fresh_policy_pairs']['accuracy']:.2f}")
10print(f"ppo_ready={ppo_ready}")1static_accuracy=0.92
2fresh_accuracy=0.64
3ppo_ready=False
KL control intuition (preview)
A common control, developed in full by the next lesson, penalizes the policy for drifting too far from its reference checkpoint. Optimization maximizes reward minus a KL-divergence term that measures policy drift.[1] That discourages large departures from the reference, but Gao's measurements are the reason to keep the claim narrow: KL control isn't a certificate that the reward model is valid on new outputs.[6] A reward model is a local approximation of human preference, which is exactly why it can be overoptimized.
Standardized evaluation: RewardBench
Held-out pairs you wrote yourself can share your blind spots. RewardBench is an Ai2 benchmark that scores a reward model by how often it ranks a known-better completion above a worse one. Its original sections cover chat, harder instruction-following comparisons, safety, reasoning, and prior preference test sets.[7]
RewardBench 2 is a harder follow-up: multi-skill, best-of-four scoring, mostly previously unused human prompts (about 70% of the set), and decontamination against twenty downstream evaluations. In their experiments, benchmark scores correlated strongly with best-of-N (Pearson 0.87 on the average they report). PPO was a weaker, saturating signal: decent-to-good reward models clustered, and transfer dropped when the reward model didn't share the policy's model lineage or prompt distribution.[8]
One practical caveat from that work: the highest-scoring reward model on the leaderboard isn't automatically the best choice for your run. Treat absolute benchmark rank as a filter, then validate with the policy and optimization setup you'll use.[8]
Your reward model tops the RewardBench leaderboard. Is it automatically the right choice for your PPO run?
Answer
No. A high benchmark score is a good filter, especially for best-of-N. RewardBench 2 reported better PPO transfer for same-lineage, in-distribution reward and policy models, and PPO gains saturated across many decent scores. Validate with your intended policy and optimization setup.
Other ways to build a preference signal
The scalar Bradley-Terry head is a common baseline. It isn't the only judge.
| Signal | What it outputs | When it helps | Main failure |
|---|---|---|---|
| Scalar Bradley-Terry RM | One number per completion | PPO, best-of-N, inspectable scores | Overoptimization, length and style shortcuts |
| DPO | Implicit reward inside the policy | Clean offline preferences, no separate judge | No reusable scalar for scoring or online RL |
| Process reward model (PRM) | A score per reasoning step | Multi-step math or tool traces | Expensive step labels; still gameable |
| Generative RM / LLM-as-judge | A verdict, optionally with a rationale | When chain-of-thought or vote aggregation helps in your eval | Cost, position bias, and judge-specific quirks |
| Verifier / RLVR | A checkable pass/fail | Math answers, tests, other exact graders | Misspecified tests and verifier gaming |
- Generative reward models. Instead of a scalar head, an LM can read candidates and emit a verdict. Mahan et al. study rationale generation and vote aggregation in that setup; treat those as design choices to evaluate, not universal guarantees.[9]
- Process reward models. Scoring only the final answer is a weak signal for multi-step reasoning. PRMs score each step. Lightman et al. compared process and outcome supervision as search over many sampled MATH solutions (best-of-N), not as RL on the generator, and the process RM beat the outcome RM.[10]
- Verifiers and RLVR. When correctness is checkable, such as a final math answer or passing unit tests, verifiable rewards can reduce dependence on a learned preference proxy. They don't eliminate misspecified tests or gaming of the verifier.[11] Later chapters cover this family.
These don't retire the scalar reward model, and they don't all share its Bradley-Terry objective. The reusable lesson is the evaluation discipline: inspect the signal's coverage, test it on outputs produced by the system being optimized, and watch for optimization exploiting its blind spots.
When an explicit reward model is worth the cost
Use one when:
- you want PPO-style online optimization
- you want to score many candidate outputs with one scalar model
- you want to inspect and audit the preference signal directly
Start with DPO when:
- you have a clean offline preference dataset
- you want the simpler baseline first
- you don't need an explicit learned judge in the loop
That trade-off is why DPO is a strong offline-preference baseline: it removes the separate reward-model training stage.[2] It doesn't supply a reusable scalar judge for PPO or candidate scoring. The final toy gate combines several checks, with project-specific thresholds, to block optimization when only one static slice passes.
1checks = {
2 "grouped_split_has_no_prompt_overlap": True,
3 "length_matched_accuracy": 0.84,
4 "fresh_policy_human_accuracy": 0.78,
5 "minimum_required_accuracy": 0.80,
6}
7ready = (
8 checks["grouped_split_has_no_prompt_overlap"]
9 and checks["length_matched_accuracy"] >= checks["minimum_required_accuracy"]
10 and checks["fresh_policy_human_accuracy"] >= checks["minimum_required_accuracy"]
11)
12
13print(f"static_slice_passes={checks['length_matched_accuracy'] >= checks['minimum_required_accuracy']}")
14print(f"fresh_policy_passes={checks['fresh_policy_human_accuracy'] >= checks['minimum_required_accuracy']}")
15print(f"optimize_against_reward={ready}")1static_slice_passes=True
2fresh_policy_passes=False
3optimize_against_reward=FalseYour reward model has good static pair accuracy, but once PPO starts, reward climbs while human raters say answers are getting verbose and manipulative. What is the first diagnosis?
Answer
The reward model is being exploited under distribution shift. Static pair accuracy wasn't enough to prove that it would rank fresh policy outputs the way humans do.
You have a clean offline preference dataset and no need to score large candidate pools or run PPO. What simpler baseline should you evaluate first?
Answer
DPO. If you don't need an explicit learned scalar judge in the loop, it removes the separate reward-model stage.
Common pitfalls
Symptom: loss falls but held-out rankings barely improve
- Cause: chosen and rejected responses are nearly equivalent, so the pair gives little ranking signal.
- Fix: audit pair strength before training. Keep pairs where the preference is clear, policy-relevant, and tied to the same prompt.
Symptom: one labeler style dominates the reward model
- Cause: inconsistent or narrow annotator preferences become inconsistent rewards.
- Fix: measure agreement, review disagreements, and separate policy rules from personal style before training.
Symptom: PPO reward rises while human preference gets worse
- Cause: the policy has found a shortcut in the reward model under distribution shift.
- Fix: add fresh policy-output evaluations, length-matched checks, adversarial probes, and human preference gates before trusting the scalar reward.
Symptom: product dashboards treat reward as truth
- Cause: reward is being mistaken for the business or human objective itself.
- Fix: report reward beside held-out preference, refusal quality, helpfulness, safety, and downstream product metrics.