Personalize this lesson
Adapt explanations and teaching visuals to your background and preferred voice.
A pairwise reward model learns to score completed responses from preference comparisons. Unlike supervised fine-tuning (SFT), which trains a network to imitate demonstration tokens, reward modeling fits an evaluative judge that separates preferred answers from rejected ones.
Consider an access-policy assistant in an enterprise setting. Company policy P-7 requires an approved access-review ticket before granting production database credentials. A user submits an urgent request during a database migration. One completion follows rule P-7 by opening a ticket and citing policy; a second grants admin access immediately to appear helpful. A third follows the rule but pads the response with three paragraphs of apologetic disclaimers. If the evaluative judge rewards raw length or deferential tone rather than policy compliance, it'll favor the wrong response. We'll use this running access-policy example to separate learning genuine human preference from picking up superficial training shortcuts.
System diagrams for RLHF (Reinforcement Learning from Human Feedback) often compress this pipeline into a box labeled reward model, before jumping to Proximal Policy Optimization (PPO). In the scalar-model recipe, that box is a neural network trained on noisy comparisons.[1] If it learns a shortcut such as verbosity, policy optimization can amplify the mistake.[2] We'll build the judge, inspect its gradients, and try to fool it before using it to train a generator.
A scalar judge, not a generator
This lesson trains a scalar reward model. Generative judges also exist, but the head we build here returns a score. An autoregressive language model instead computes a conditional probability distribution over the vocabulary at each step:
A reward model, by contrast, acts as a sequence scorer. It ingests a completed sequence , processes all tokens through its transformer backbone, and projects the final representation down to a single real-valued scalar:
For prompt , suppose an annotator evaluates two completions:
- (chosen / winner): opens an access-review ticket and cites policy P-7.
- (rejected / loser): grants temporary admin credentials immediately.
The model's objective is to assign scores satisfying the preference inequality:
1r(x, y_w) > r(x, y_l)Pairwise comparisons supervise relative ranking, not an absolute metric scale. A score of 2.0 isn't twice as good as 1.0, and raw scores from two independently trained reward models can't be compared directly without shared calibration.
That scalar score serves two major downstream roles in production post-training:
- Terminal reward for online RL: The scorer evaluates completed policy rollouts. PPO commonly combines those rewards with a learned value baseline to compute Generalized Advantage Estimation (GAE). Outcome-supervised GRPO instead computes advantages relative to a group of responses to the same prompt, without that value model.[3]
- Best-of-N selection: During inference or synthetic data generation, sample candidate completions and return . This is often called rejection sampling in post-training recipes; it isn't exact sampling from a specified target distribution.
Direct Preference Optimization (DPO) uses the analytic relationship between a reward and its optimal policy under KL-regularized reward maximization. Reparameterizing the reward as a policy/reference log-ratio turns the Bradley-Terry preference loss into a policy loss.[4] The basic offline recipe needs labeled pairs, without a separate scorer. An explicit reward model is useful when you want a reusable, fast ranking signal. Online exploration, best-of-N selection, and reasoning verification can also use a generative judge, human labels, or an executable verifier; none inherently requires a learned scalar head.[5]
Preference data anatomy and validation pipeline
Check what the preference data actually labels before fitting the scorer. The loss rewards the supplied comparisons, including any systematic annotation mistakes.
Standard format
An explicit-prompt row stores the shared user query in prompt and isolates the competing responses in chosen and rejected.
1{
2 "prompt": "Policy P-7 requires an approved access-review ticket before granting admin access. A user requests temporary admin access for a migration. What should the assistant do?",
3 "chosen": "Open an access-review ticket and cite policy P-7 before approval.",
4 "rejected": "Grant admin for tonight and ask the user to clean it up tomorrow."
5}Conversational format
Multi-turn dialogues follow the same pairwise contract using structured role dictionaries. Both candidates share identical system instructions and user history, branching only at the final assistant turn.
1{
2 "chosen": [
3 {"role": "system", "content": "Policy P-7 requires an approved access-review ticket before granting admin access."},
4 {"role": "user", "content": "User requests temporary admin access for a migration. What should the assistant do?"},
5 {"role": "assistant", "content": "Open an access-review ticket and cite policy P-7 before approval."}
6 ],
7 "rejected": [
8 {"role": "system", "content": "Policy P-7 requires an approved access-review ticket before granting admin access."},
9 {"role": "user", "content": "User requests temporary admin access for a migration. What should the assistant do?"},
10 {"role": "assistant", "content": "Grant admin for tonight and ask the user to clean it up tomorrow."}
11 ]
12}Hugging Face TRL's RewardTrainer supports standard strings and conversational message lists, with implicit or explicit prompts. Conversational inputs use the tokenizer's chat template; standard strings are concatenated and tokenized without automatically becoming a chat conversation. These behaviors were checked against TRL 1.13.0 on September 22, 2026.[6]
A pair contract before training
A dataset isn't ready for training simply because it has chosen and rejected fields. Enforce this data contract before feeding pairs into your training script:
- Identical context prefix: Both candidates must share the exact prompt, system prompt, tool definitions, and formatting template.
- Distinct scored inputs: If both candidates become identical token IDs and masks, a deterministic shared scorer produces zero score difference and zero pair-loss parameter gradient. Check after templating and tokenization as well as before them.
- Labels matched to the objective: The winner-only loss below expects a strict preference. Route ties and abstentions separately. Low agreement can indicate a bad rubric, an error, or legitimate differences in taste; review it rather than automatically deleting it. Repeated votes can supervise a soft preference probability.
- Lineage metadata: Track the generating model checkpoint, decoding temperature, and annotator cohort for every candidate.
This small audit admits only strict chosen-over-rejected labels. It catches text-level problems; a real pipeline also checks tokenized identity. Its tie filter is a choice for this binary exercise, not a rule for every preference dataset.
1from collections import Counter
2
3pairs = [
4 {"id": "a", "prompt_left": "admin?", "prompt_right": "admin?", "chosen": "Escalate.", "rejected": "Approve.", "label": "chosen"},
5 {"id": "b", "prompt_left": "admin?", "prompt_right": "admin?", "chosen": "Escalate.", "rejected": "Escalate.", "label": "chosen"},
6 {"id": "c", "prompt_left": "admin?", "prompt_right": "source?", "chosen": "Escalate.", "rejected": "Cite source.", "label": "chosen"},
7 {"id": "d", "prompt_left": "source?", "prompt_right": "source?", "chosen": "Cite.", "rejected": "Refuse.", "label": "tie"},
8]
9
10def rejection_reason(pair):
11 fields = ("prompt_left", "prompt_right", "chosen", "rejected")
12 if any(not isinstance(pair.get(field), str) or not pair[field].strip() for field in fields):
13 return "missing_or_empty_text"
14 if pair.get("label") != "chosen":
15 return "tie_or_abstention"
16 if pair["prompt_left"] != pair["prompt_right"]:
17 return "context_mismatch"
18 if pair["chosen"].strip() == pair["rejected"].strip():
19 return "identical_candidates"
20 return None
21
22reasons = Counter(reason for pair in pairs if (reason := rejection_reason(pair)))
23kept = [pair["id"] for pair in pairs if rejection_reason(pair) is None]
24
25print(f"kept={kept}")
26print(f"rejected={dict(sorted(reasons.items()))}")
27assert kept == ["a"]
28assert rejection_reason({**pairs[0], "chosen": " "}) == "missing_or_empty_text"
29assert rejection_reason({}) == "missing_or_empty_text"1kept=['a']
2rejected={'context_mismatch': 1, 'identical_candidates': 1, 'tie_or_abstention': 1}Decide what the evaluation split should measure
A prompt with four candidate completions yields six pairwise comparisons. Splitting those pairs independently can put the same prompt and even the same completion into training and evaluation. That overlap makes the evaluation a poor test of generalization to unseen prompts.
For unseen-prompt generalization, group by prompt ID and inspect near-duplicate prompt clusters. Keep a group's comparisons together. If the intended task is ranking new responses to a fixed set of known prompts, a response-held-out split can answer that different question. Describe the split and prevent unintended reuse of evaluation completions or labels.
1pairs = [
2 {"pair_id": "a-b", "prompt_id": "support-17"},
3 {"pair_id": "a-c", "prompt_id": "support-17"},
4 {"pair_id": "d-e", "prompt_id": "safety-04"},
5 {"pair_id": "f-g", "prompt_id": "access-09"},
6]
7eval_prompt_ids = {"support-17"}
8
9train = [pair for pair in pairs if pair["prompt_id"] not in eval_prompt_ids]
10evaluation = [pair for pair in pairs if pair["prompt_id"] in eval_prompt_ids]
11overlap = {pair["prompt_id"] for pair in train} & {pair["prompt_id"] for pair in evaluation}
12
13assert not overlap
14print(f"train_pairs={len(train)} eval_pairs={len(evaluation)}")
15print(f"prompt_overlap={sorted(overlap)}")1train_pairs=2 eval_pairs=2
2prompt_overlap=[]InstructGPT asked labelers to rank to completions per prompt, producing correlated comparisons. Its authors observed overfitting when those comparisons were shuffled as separate examples, and instead grouped them in a single batch item.[1] Each completion participates in pairs; whether those appearances cause separate optimizer updates depends on batching.
Grouping also lets you score each of the completions once and reuse those scores across comparisons. Uniformly sampling prompt groups and averaging each group's pair losses avoids giving a prompt 36 times the weight of a prompt merely because it has more pairs. It equalizes the prompt's loss weight, not the norm of every prompt's gradient.
Give the head a representation of the whole answer
How do you turn a generative transformer into a preference judge?
Choose a backbone and replace the vocabulary head
An SFT checkpoint is a common starting point, as in InstructGPT.[1] It is not mandatory: Llama 3 trained its reward model on a pretrained checkpoint.[7] Choose a backbone with suitable capacity, language coverage, context length, and formatting, then evaluate it on the generator's actual outputs.
The architectural surgery is straightforward:
- Strip the language modeling unembedding projection layer ().
- Add a scalar linear head (), optionally with a bias. A shared scalar bias cancels from pairwise differences and receives no gradient from the pair loss alone.
- Train the backbone, train LoRA adapters plus the head, or freeze the backbone and train only the head. With adapters, ensure the head is trainable and included in the saved artifact.[6]
1Input Tokens -> Transformer Hidden States (d) -> Last-Token Representation -> Linear Head -> Scalar Reward r(x, y)Why the final real token is a useful readout
In ordinary causal decoder attention, position may attend to positions , including itself. The allowed attention pattern is lower triangular including the diagonal.
The final non-padding token, often a template's end-of-turn token, can incorporate the full preceding prompt and response. An earlier token cannot see the response suffix after its position. A dedicated scoring token appended after the answer can work too, and other architectures can use other pooling schemes. The requirement is to give the scorer the answer it is supposed to judge.
For conventional contiguous padding:
- Left-padding: A nonempty sequence's final real token sits at index
-1. - Right-padding:
last_idx = mask.sum(dim=-1) - 1works when real tokens start at index zero and there are no internal mask gaps.
Do not assume a hidden state at a padding position is a clean padding embedding. Its attention behavior depends on the model. Gather the intended readout explicitly and check your sequence-classification model's pooling implementation.
Sequence length filtering vs truncation hazards
Truncation can change what a label means. If the preference hinged on a factual error at the end, removing that suffix can remove the evidence for the label. Filtering long pairs avoids that mismatch, but also changes the training distribution; report how many pairs and which domains are removed.
In TRL 1.13.0, RewardConfig.max_length defaults to 1024 and filters a pair if either tokenized candidate exceeds it; None disables that filtering. Count the full prompt-plus-response sequence, including template and end tokens.[6] These toy lengths illustrate the boundary without loading a tokenizer.
1max_length = 1024
2pairs = [
3 {"id": "fits", "chosen_tokens": 412, "rejected_tokens": 390},
4 {"id": "boundary", "chosen_tokens": 1024, "rejected_tokens": 1024},
5 {"id": "chosen_too_long", "chosen_tokens": 1088, "rejected_tokens": 401},
6 {"id": "rejected_too_long", "chosen_tokens": 288, "rejected_tokens": 1030},
7]
8
9kept = [p["id"] for p in pairs if max(p["chosen_tokens"], p["rejected_tokens"]) <= max_length]
10dropped = [p["id"] for p in pairs if p["id"] not in kept]
11
12print(f"kept={kept}")
13print(f"dropped_instead_of_truncated={dropped}")
14assert kept == ["fits", "boundary"]1kept=['fits', 'boundary']
2dropped_instead_of_truncated=['chosen_too_long', 'rejected_too_long']The script below shows how to correctly pool the last non-padding token representation regardless of whether padding appears on the left or right.
1import numpy as np
2
3def last_nonpadding(hidden, mask):
4 hidden, mask = np.asarray(hidden), np.asarray(mask)
5 if hidden.ndim != 3 or mask.shape != hidden.shape[:2]:
6 raise ValueError("expected hidden [batch, time, width] and matching mask")
7 if not np.isin(mask, [0, 1]).all() or not mask.any(axis=1).all():
8 raise ValueError("each sequence needs at least one unmasked position")
9 positions = np.arange(mask.shape[1])
10 last = np.where(mask == 1, positions, -1).max(axis=1)
11 return hidden[np.arange(len(hidden)), last]
12
13states = np.array([[[1., 0.], [2., 3.], [99., 99.]],
14 [[99., 99.], [1., 0.], [2., 3.]]])
15pooled = last_nonpadding(states, [[1, 1, 0], [0, 1, 1]])
16scores = pooled @ np.array([0.5, 1.0])
17assert np.allclose(scores, [4.0, 4.0])
18try:
19 last_nonpadding(states[:1], [[0, 0, 0]])
20except ValueError:
21 pass
22else:
23 raise AssertionError("empty sequence accepted")
24print("right/left-padded scores:", scores.tolist())1right/left-padded scores: [4.0, 4.0]Bradley-Terry formulation, BCE loss, and gradient dynamics
Return to the access-policy pair:
- Chosen completion (ticket and cite P-7) gets scalar reward .
- Rejected completion (grant admin tonight) gets scalar reward .
Their margin is . How does that difference translate into a probabilistic preference?

Bradley-Terry pairwise preference formulation
The Bradley-Terry (1952) model formulates the probability that candidate is preferred over candidate given prompt as:
Dividing numerator and denominator by yields the standard logistic sigmoid over the score difference:
For our access-policy margin :
A positive margin assigns the chosen completion more than 50% win probability. At zero, the model assigns equal probability to either winner. That is not an epistemic uncertainty estimate, a tie probability, or evidence that humans actually split 50-50.
Direct equivalence to binary cross-entropy with logits
Training maximizes the log-likelihood of observing the preferred outcomes across dataset :
Notice that this is mathematically identical to binary cross-entropy (BCE) where target indicates :
Use the logit difference directly in PyTorch's stable BCE implementation. This executable check also exposes the gradient on a shared score bias:
1import torch
2
3# One row per pair, columns chosen/rejected.
4rewards = torch.tensor([[1.8, 0.7]], dtype=torch.float64, requires_grad=True)
5bias = torch.tensor(10.0, dtype=torch.float64, requires_grad=True)
6scores = rewards + bias
7logits_diff = scores[:, 0] - scores[:, 1]
8loss = torch.nn.functional.binary_cross_entropy_with_logits(
9 logits_diff, torch.ones_like(logits_diff)
10)
11assert torch.allclose(loss, torch.nn.functional.softplus(-logits_diff).mean())
12loss.backward()
13expected = torch.sigmoid(logits_diff.detach()) - 1
14assert torch.allclose(rewards.grad[:, 0], expected)
15assert torch.allclose(rewards.grad[:, 1], -expected)
16assert bias.grad.item() == 0.0
17print(f"loss={loss.item():.4f}")
18print(f"chosen_gradient={rewards.grad[0, 0].item():+.4f} rejected_gradient={rewards.grad[0, 1].item():+.4f}")
19print(f"shared_bias_gradient={bias.grad.item():.1f}")1loss=0.2873
2chosen_gradient=-0.2497 rejected_gradient=+0.2497
3shared_bias_gradient=0.0Gradient dynamics and the symmetric push
Let . Differentiating with respect to :
By the chain rule, the gradients on individual rewards are:
This derivative governs model learning:
- Confident correct ranking (): The pair-loss derivatives with respect to the scores approach zero. Parameter updates also depend on the scorer's Jacobian, other loss terms, and the optimizer.
- Ambiguous ranking (): . Gradients are on and on .
- Confidently wrong ranking (): . Gradients approach maximum magnitude: pulling up, and pushing down!
The score origin is free; its difference scale is not
Because the loss depends only on the difference , adding any prompt-dependent scalar to all scores leaves the margin completely unchanged:
Scores 11.80 and 10.70 produce the same loss, win probability, and score derivatives as 1.80 and 0.70. Pairwise loss can't anchor the absolute score origin without an explicit centering constraint.
Multiplying both scores by 3 is different: it triples the margin and changes the sigmoid probability and loss. The fixed logistic link gives score differences a log-odds scale. On separable training data, a model can keep increasing margins to reduce loss, so regularization and held-out probability checks still matter.
The script below demonstrates numerically stable negative log-sigmoid evaluation across three representative margins: confident positive, incorrect negative, and mild positive.
1from math import exp, log1p
2from statistics import mean
3
4chosen_rewards = [2.1, 0.8, 1.9]
5rejected_rewards = [0.4, 1.0, 1.2]
6margins = [a - b for a, b in zip(chosen_rewards, rejected_rewards)]
7
8def neg_log_sigmoid(z: float) -> float:
9 if z >= 0:
10 return log1p(exp(-z))
11 return -z + log1p(exp(z))
12
13loss = mean(neg_log_sigmoid(m) for m in margins)
14accuracy = mean(m > 0 for m in margins)
15
16print("margins:", [round(m, 3) for m in margins])
17print("reward_loss:", round(loss, 4))
18print("pair_accuracy:", round(accuracy, 4))
19assert abs(neg_log_sigmoid(0.0) - log1p(1.0)) < 1e-12
20assert neg_log_sigmoid(-1000.0) == 1000.0
21assert neg_log_sigmoid(1000.0) == 0.0 # float underflow is harmless here1margins: [1.7, -0.2, 0.7]
2reward_loss: 0.4564
3pair_accuracy: 0.6667The script below evaluates loss across three points on the margin curve: wrong negative, neutral zero, and confident positive.
1from math import exp, log1p
2
3def neg_log_sigmoid(z: float) -> float:
4 if z >= 0:
5 return log1p(exp(-z))
6 return -z + log1p(exp(z))
7
8for margin in [-2.0, 0.0, 2.0]:
9 print(f"margin={margin:+.1f} loss={neg_log_sigmoid(margin):.4f}")1margin=-2.0 loss=2.1269
2margin=+0.0 loss=0.6931
3margin=+2.0 loss=0.1269Training a small reward head and isolating shortcut features
Let's examine how easily a reward model can learn a spurious shortcut instead of the intended policy behavior.
Fitting a linear head on CPU
We train a linear reward head on two hand-assigned features: policy compliance and scaled response length.
In the confounded training set, compliant answers are always one length unit longer than non-compliant answers. In the balanced training set, short compliant and long non-compliant pairs break that correlation.
Both heads achieve 100% training accuracy. But when evaluated on counterexamples where the non-compliant answer is verbose, the confounded head collapses to 33% accuracy!
1import json
2import tempfile
3from pathlib import Path
4import numpy as np
5
6def objective_and_gradient(weights, pairs, l2=0.02):
7 pairs = np.asarray(pairs, dtype=float)
8 if pairs.ndim != 3 or pairs.shape[1:] != (2, 2) or len(pairs) == 0:
9 raise ValueError("expected nonempty [pairs, chosen/rejected, two features]")
10 if not np.isfinite(pairs).all():
11 raise ValueError("features must be finite")
12 differences = pairs[:, 0] - pairs[:, 1]
13 margins = differences @ weights
14 loss = np.logaddexp(0.0, -margins).mean() + 0.5 * l2 * (weights @ weights)
15 sigmoid_negative = np.exp(-np.logaddexp(0.0, margins))
16 gradient = -(differences.T @ sigmoid_negative) / len(pairs) + l2 * weights
17 return float(loss), gradient
18
19def fit(pairs):
20 weights = np.zeros(2)
21 initial, _ = objective_and_gradient(weights, pairs)
22 for _ in range(300):
23 _, gradient = objective_and_gradient(weights, pairs)
24 weights -= 0.1 * gradient
25 final, _ = objective_and_gradient(weights, pairs)
26 assert final < initial
27 return weights
28
29def accuracy(weights, pairs):
30 return float(np.mean((pairs[:, 0] - pairs[:, 1]) @ weights > 0))
31
32# Each row is [chosen representation, rejected representation].
33confounded = np.array([[[1., 1.], [0., 0.]], [[1., 2.], [0., 1.]]])
34balanced = np.array([[[1., 1.], [0., 0.]], [[1., 0.], [0., 1.]]])
35counterexamples = np.array([[[1., 0.], [0., 2.]],
36 [[1., 0.], [0., 3.]],
37 [[1., 1.], [0., 1.]]])
38
39# Check both analytic gradients against central finite differences.
40probe = np.array([0.2, -0.3])
41_, analytic = objective_and_gradient(probe, confounded)
42for invalid in [np.empty((0, 2, 2)), np.zeros((2, 2)), np.full((1, 2, 2), np.nan)]:
43 try:
44 objective_and_gradient(probe, invalid)
45 raise AssertionError("invalid pair tensor accepted")
46 except ValueError:
47 pass
48for coordinate in range(2):
49 delta = np.zeros(2)
50 delta[coordinate] = 1e-6
51 numerical = (objective_and_gradient(probe + delta, confounded)[0]
52 - objective_and_gradient(probe - delta, confounded)[0]) / 2e-6
53 assert abs(numerical - analytic[coordinate]) < 1e-7
54
55for name, training in [("confounded", confounded), ("balanced", balanced)]:
56 weights = fit(training)
57 print(f"{name}: weights={weights.round(3).tolist()}, train={accuracy(weights, training):.2f}, counterexamples={accuracy(weights, counterexamples):.2f}")
58 if name == "confounded":
59 assert accuracy(weights, counterexamples) == 1 / 3
60 else:
61 assert accuracy(weights, counterexamples) == 1.0
62
63# Save the fitted head together with its feature order; reload without refitting.
64with tempfile.TemporaryDirectory() as directory:
65 path = Path(directory) / "reward_head.json"
66 path.write_text(json.dumps({"features": ["compliance", "scaled_length"], "weights": weights.tolist()}))
67 artifact = json.loads(path.read_text())
68 assert artifact["features"] == ["compliance", "scaled_length"]
69 assert np.allclose(counterexamples @ np.array(artifact["weights"]), counterexamples @ weights)
70print("gradient check and head reload: passed")1confounded: weights=[1.629, 1.629], train=1.00, counterexamples=0.33
2balanced: weights=[2.667, 0.0], train=1.00, counterexamples=1.00
3gradient check and head reload: passed![Measured results from the two-feature CPU experiment: both heads achieve 100% training accuracy. On three length-controlled counterexamples, the confounded head (weights [1.63, 1.63]) drops to 33% accuracy by following a length shortcut, while the balanced head (weights [2.67, 0.00]) achieves 100% accuracy.](/cdn/content-image/training/reward-modeling-from-preference-data/illustrations/_generated/reward_generalization_gap_dark.png?v=8681ff4bea93)
After 300 updates, the confounded head has . Symmetric initialization and identical compliance/length differences give both features equal weight. On the first counterexample, a verbose violation scores , beating the compliant answer's 1.629.
The balanced set has difference vectors and . Starting at zero length weight, their length gradients cancel, keeping that weight at zero. L2 regularization shrinks weights; it does not by itself identify length as the unwanted feature. After 300 updates, passes these three counterexamples. This is a feature-constructed CPU experiment, not a transformer training run or a broad robustness result.
Suppose you remove L2 regularization but keep the balanced pairs and zero initialization. Does the length weight start growing?
Answer
No. The two pairs have the same compliance difference and opposite length differences. Their length gradients still cancel at zero length weight. This isolates the effect of the data construction from the effect of regularization.
Ranking expansions and margin-augmented losses
For a ranked group of completions, two common modeling choices are:
- Pair expansion: Generate all pairs and average their loss within the prompt group.[1]
- Plackett-Luce formulation: Model the full permutation as a multi-class choice at each rank.
Some pipelines record degree of preference, such as "much better" or "slightly better." Llama 2 used a rating-dependent margin inside the sigmoid; the weakest category had margin zero.[8]
A margin shifts the loss curve. At , the effective logit is zero, the loss is , and its margin derivative is ; the loss has not saturated. Achieving effective win probability requires . For and , the required difference is about 3.197. Llama 3 removed the margin after observing diminishing improvements with data scaling.[7]
1from math import exp, log, log1p, isclose
2
3m = 1.0
4required_difference = m + log(0.9 / 0.1)
5for difference in [m, required_difference]:
6 logit = difference - m
7 probability = 1 / (1 + exp(-logit))
8 loss = log1p(exp(-logit))
9 print(f"score_difference={difference:.3f} effective_p={probability:.2f} loss={loss:.4f} derivative={probability - 1:+.2f}")
10assert isclose(1 / (1 + exp(-(required_difference - m))), 0.9)1score_difference=1.000 effective_p=0.50 loss=0.6931 derivative=-0.50
2score_difference=3.197 effective_p=0.90 loss=0.1054 derivative=-0.10Metrics, centering invariance, and calibration
Tracking pair accuracy alone during reward model training is a notorious trap.
Why pair accuracy alone is insufficient
Pair accuracy is a discrete indicator: . It hides critical failure modes that degrade downstream RL:
| Logged Metric | What It Measures | What It Conceals |
|---|---|---|
| Pair accuracy | Binary win frequency () | Collapse of margin magnitudes or shortcut reliance |
| Mean margin | Separation distance | Score calibration and probability alignment |
| Mean reward | Overall score drift | Relative ordering quality across prompts |
| Gradient norm | Optimization stability | Reliance on spurious features such as length |
| Fresh-policy pair accuracy | Agreement on newly generated policy text | Unsampled failures, probability calibration, and human disagreement |
Centering invariance and PPO scale sensitivity
In PPO, the policy optimization objective balances raw reward against a KL divergence penalty:
Multiplying rewards by 3 while keeping fixed gives the same ideal maximizer as using the original rewards with . The relative KL penalty becomes weaker. That does not guarantee uncontrolled drift: clipping, optimization budgets, normalization, and the reward landscape affect the actual run. Subtracting a fixed prompt-dependent offset changes neither candidate ranking nor this ideal maximizer; it does not correct a scale mismatch.
InstructGPT shifted scores so labeler demonstrations had mean zero before RL; this anchored the origin.[1] TRL 1.13.0 exposes center_rewards_coefficient. Its documentation recommends trying 0.01, while the default is None (disabled). The actual auxiliary term is:[6]
Writing the pair midpoint as gives . There is no explicit squared-margin term: scores and have zero centering penalty. This encourages pair midpoints near zero; it neither bounds every score nor calibrates probabilities. Joint training through shared weights can still change margins. This regularizer comes from Eisenstein et al.'s ensemble study.[9]
The script below demonstrates that shifting both scores by a constant leaves pair loss strictly invariant, while scaling scores alters the loss slope.
1from math import exp, log1p
2from statistics import mean
3
4chosen = [1.2, 0.7]
5rejected = [0.2, 0.4]
6
7def pair_loss(left, right):
8 return mean(log1p(exp(-(a - b))) for a, b in zip(left, right))
9
10shifted = ([score + 10 for score in chosen], [score + 10 for score in rejected])
11scaled = ([score * 3 for score in chosen], [score * 3 for score in rejected])
12
13print(f"base_loss={pair_loss(chosen, rejected):.4f}")
14print(f"shifted_loss={pair_loss(*shifted):.4f}")
15print(f"scaled_loss={pair_loss(*scaled):.4f}")
16gamma = 0.01
17sum_penalty = gamma * (100 + (-100)) ** 2
18individual_penalty = gamma * (100 ** 2 + (-100) ** 2) / 2
19assert sum_penalty == 0.0 and individual_penalty == 100.0
20print(f"scores=+100,-100 pair_sum_penalty={sum_penalty:.1f} individual_square_penalty={individual_penalty:.1f}")1base_loss=0.4338
2shifted_loss=0.4338
3scaled_loss=0.1949
4scores=+100,-100 pair_sum_penalty=0.0 individual_square_penalty=100.0Probabilistic calibration and Brier scores
A reward difference gives an implied win probability, not a measured agreement rate. To evaluate calibration, fix candidate labels A and B before observing the human vote. Let when A wins and when B wins, and . Do not reorder every row so the observed winner is A and then treat all targets as 1.
The Brier score measures probability error:
For a binary-probability version of Expected Calibration Error (ECE), bin and compare each bin's mean prediction with its observed A-win frequency:
Another common ECE reports predicted-class confidence versus classification accuracy; state which convention you use. ECE depends on the binning and sample size. Report counts and evaluate slices as well as the aggregate.
Imagine four independent raters compare the same A/B pair. Three choose A and one chooses B. A model predicting 0.75 and one predicting 0.99 have the same winner predictions. Which better reflects these observed votes?
1from statistics import mean
2
3# A/B order was fixed before annotation. This is a tiny illustrative sample.
4votes = [1, 1, 1, 0]
5for probability in [0.75, 0.99]:
6 accuracy = mean(int(probability > 0.5) == vote for vote in votes)
7 brier = mean((probability - vote) ** 2 for vote in votes)
8 # All identical predictions occupy one bin.
9 one_bin_ece = abs(probability - mean(votes))
10 assert accuracy == 0.75
11 print(f"p_A={probability:.2f} accuracy={accuracy:.2f} Brier={brier:.4f} one_bin_ECE={one_bin_ece:.2f}")1p_A=0.75 accuracy=0.75 Brier=0.1875 one_bin_ECE=0.00
2p_A=0.99 accuracy=0.75 Brier=0.2451 one_bin_ECE=0.24The 0.75 prediction better matches this sample. Four votes cannot establish population calibration. For training with repeated votes, a soft target gives loss , minimized at when . A 50-50 vote split and an explicit "equally good" annotation are different observations; a tie-aware preference model may be appropriate for the latter.
Reward hacking, Goodhart's Law, and policy overoptimization
How can a reward model with high static validation accuracy still lead a policy toward poor outputs?
The Goodhart mechanism and Gao et al. scaling laws
Goodhart's Law describes the risk of optimizing a proxy until it stops tracking the intended goal. A measure does not inevitably fail merely because it is used as a target.
In RLHF, is a proxy for the preference criteria we intend to optimize. Gao et al. (2023) studied overoptimization using a larger gold reward model, not direct human ratings, to generate labels and evaluate policies. In their setup, more optimization could raise proxy reward while gold-model reward peaked and then declined. They fitted empirical scaling relationships for PPO and best-of-N; these are not universal laws of human satisfaction or evidence that proxy scores grow to infinity.[2]
The mechanism is useful: optimization selects outputs with high proxy scores, including outputs scored highly because of evaluator error. As the policy's output distribution changes, errors that were rare in a static split can matter much more. Monitor independent evaluation during optimization rather than extrapolating static accuracy.
Exploitation vectors: verbosity, sycophancy, and format bias
When an RL optimizer searches through completion space, it exploits specific failure modes in the reward proxy:
- Verbosity bias: Length can correlate with preferred answers in the training data, letting the scorer favor irrelevant elaboration. Longer answers are sometimes better; test whether added length helps when useful content is held fixed.
- Sycophancy: A scorer can favor agreement with the user's mistaken premise over a truthful correction. Sharma et al. found evidence of this in human preferences and preference models; politeness or an apology alone is not sycophancy.[10]
- Formatting shortcuts: Structure can improve readability, but a scorer may reward familiar formatting even when the answer is wrong. Compare responses whose factual content and presentation vary independently.[9]
- Misspecified criteria: A score trained to favor helpfulness may omit privacy or factuality. A scalar can encode a well-specified utility; its dimensionality alone does not prove that hacking is inevitable. Inspect what the labels and reward function actually reward.
This audit flags a possible length confound. Three chosen answers being longer is not proof that the scorer uses length; the controlled counterexamples above test that mechanism.
1from statistics import mean
2
3pairs = [
4 {"chosen_reward": 1.8, "rejected_reward": 0.7, "chosen_tokens": 36, "rejected_tokens": 19},
5 {"chosen_reward": 2.4, "rejected_reward": 1.1, "chosen_tokens": 58, "rejected_tokens": 22},
6 {"chosen_reward": 1.6, "rejected_reward": 0.3, "chosen_tokens": 33, "rejected_tokens": 14},
7]
8
9margins = [row["chosen_reward"] - row["rejected_reward"] for row in pairs]
10accuracy = mean(margin > 0 for margin in margins)
11length_gaps = [row["chosen_tokens"] - row["rejected_tokens"] for row in pairs]
12
13print(f"pair_accuracy={accuracy:.2f}")
14print(f"mean_margin={mean(margins):.2f}")
15print(f"chosen_answers_longer={all(gap > 0 for gap in length_gaps)}")
16print("next_check=build length-matched eval pairs")1pair_accuracy=1.00
2mean_margin=1.23
3chosen_answers_longer=True
4next_check=build length-matched eval pairsThis triage sends disagreements for review. It deliberately uses unanimity as a small exercise's threshold. A real annotation policy may retain disputed pairs, collect more votes, or model cohort preferences; unanimity is not a universal training requirement.
1votes = {
2 "clear_safety": ["chosen", "chosen", "chosen"],
3 "style_only": ["chosen", "rejected", "chosen"],
4 "ambiguous_refusal": ["chosen", "rejected", "tie"],
5}
6
7for pair_id, labels in votes.items():
8 chosen_share = labels.count("chosen") / len(labels)
9 status = "train" if chosen_share == 1.0 else "review_or_hold_out"
10 print(f"{pair_id}: chosen_share={chosen_share:.2f} status={status}")1clear_safety: chosen_share=1.00 status=train
2style_only: chosen_share=0.67 status=review_or_hold_out
3ambiguous_refusal: chosen_share=0.33 status=review_or_hold_outTest a mitigation before trusting it
1. A length margin can reward the very shortcut you want to remove
Consider this proposed margin:
Subtracting that margin makes a longer winner require a larger score difference. It does not prevent the head from scoring length. For a length-only scorer and a positive length gap, the effective logit is . Increasing still drives the loss down while preserving the shortcut.
1from math import exp, log1p
2
3alpha = 1.0
4chosen_length, rejected_length = 2.0, 1.0
5for length_weight in [0.0, 3.0, 6.0]:
6 gap = chosen_length - rejected_length
7 effective_logit = (length_weight - alpha) * gap
8 loss = max(-effective_logit, 0) + log1p(exp(-abs(effective_logit)))
9 # A different pair has a short good answer and a long wrong answer.
10 short_good_score = length_weight * 1.0
11 long_wrong_score = length_weight * 2.0
12 passes_reversal = short_good_score > long_wrong_score
13 print(f"length_weight={length_weight:.0f} training_loss={loss:.4f} passes_length_reversal={passes_reversal}")1length_weight=0 training_loss=1.3133 passes_length_reversal=False
2length_weight=3 training_loss=0.1269 passes_length_reversal=False
3length_weight=6 training_loss=0.0067 passes_length_reversal=FalseAs the length weight increases, loss improves while the long wrong answer wins the reversed pair. Start with controlled data: useful long answers, useful short answers, verbose errors, and necessary detailed explanations. Evaluate length-matched and length-reversed slices. A length correction or auxiliary objective needs its own validation; automatically penalizing all long answers can punish completeness.
2. Ensemble uncertainty quantification ()
An ensemble can expose disagreement across scorers with different seeds, data subsets, or backbones. Its errors need not be independent. Before combining raw scores, align their origins and scales on shared calibration data; otherwise arbitrary offsets can masquerade as uncertainty.[9]
During online policy optimization, compute the ensemble mean and standard deviation:
One pessimistic aggregation rule subtracts a multiple of disagreement:
This is often called a Lower Confidence Bound (LCB), but the formula alone is not a statistically calibrated confidence bound. If the scorers disagree, the penalty lowers that output's score. Unfamiliar outputs do not necessarily produce disagreement, and familiar ones can. Shared errors can receive high reward with low variance. Eisenstein et al. found that ensembles mitigated hacking but did not eliminate it.[9]
Three reward models score an answer 5, 5, and 5. The answer cites a nonexistent source. With , what does the ensemble rule return, and does it catch the error?
Answer
The mean is 5 and the standard deviation is 0, so the aggregate remains 5. Agreement cannot detect a factual error all three models share. You need evidence checks or another evaluator that can identify the missing source.
Auditing fresh policy rollouts before optimization
Sample fresh completions from the generator being aligned, construct preference pairs, and obtain independent evaluation. Include factuality, refusal quality, and controlled shortcut tests. Report sample counts, uncertainty, and disagreement. A judge ensemble needs its own audit; agreement among its members is not a substitute for ground truth.
The next example is a hypothetical screening rule. Its 80% threshold is authored for the exercise, not a research-backed PPO readiness standard. Passing it would justify further evaluation, not automatic release.
1evaluation = {
2 "static_held_out_pairs": {"accuracy": 0.92, "human_reviewed": False},
3 "fresh_policy_pairs": {"accuracy": 0.64, "human_reviewed": True},
4}
5minimum_fresh_accuracy = 0.80
6def passes_fresh_gate(result, minimum):
7 return result["human_reviewed"] and result["accuracy"] >= minimum
8
9assert not passes_fresh_gate({"accuracy": 0.99, "human_reviewed": False}, 0.80)
10assert passes_fresh_gate({"accuracy": 0.80, "human_reviewed": True}, 0.80)
11passes_pair_screen = passes_fresh_gate(evaluation["fresh_policy_pairs"], minimum_fresh_accuracy)
12
13print(f"static_accuracy={evaluation['static_held_out_pairs']['accuracy']:.2f}")
14print(f"fresh_accuracy={evaluation['fresh_policy_pairs']['accuracy']:.2f}")
15print(f"passes_pair_screen={passes_pair_screen}")1static_accuracy=0.92
2fresh_accuracy=0.64
3passes_pair_screen=False
Standardized evaluation: RewardBench and RewardBench 2
Internal pairs can share your pipeline's blind spots. Standardized benchmarks add tests from other sources, while still having their own coverage limits.
RewardBench (Lambert et al., 2024) systematically scores reward models across four core categories:[11]
- Chat (standard dialogue and helpfulness)
- Chat Hard (challenging instruction-following edge cases)
- Safety (refusals, jailbreak attempts, and harm mitigation)
- Reasoning (mathematics and code validation)
RewardBench 2 (Malik et al., 2025; revised April 2026) covers Factuality, Precise Instruction Following, Math, Safety, Focus, and Ties. Ordinary subsets select one winner among four responses, giving a 25% random baseline; Ties uses a separate metric for equivalent correct answers. Its scores are not directly comparable to the original pairwise benchmark.[12]
In the authors' Tulu 3 8B SFT experiments, benchmark scores correlated with best-of-16 downstream results (Pearson 0.87). PPO outcomes were much less sensitive to differences among decent reward models, and worsened with mismatched reward-model initialization or training-prompt distributions. These are results from that setup, not a universal ranking of PPO recipes.[12] Record benchmark revision, scorer checkpoint, tokenizer, template, and evaluation settings so your comparisons are reproducible.
Your reward model tops the RewardBench leaderboard. Can you safely skip fresh rollout validation and deploy it directly into your PPO loop?
Answer
No. Its authors found that benchmark rank did not reliably distinguish PPO outcomes in their experiments. Evaluate the scorer on your generator and target prompts, then monitor a limited optimization trial with independent checks.
Architectural spectrum: ORMs, PRMs, verifiers, and generative judges
A scalar Bradley-Terry Outcome Reward Model (ORM) isn't the only preference mechanism. Post-training architectures select from a spectrum of supervision signals:
| Approach | Feedback or objective | Primary strengths | Primary vulnerabilities |
|---|---|---|---|
| Scalar Bradley-Terry ORM | Single scalar | Fast scoring, PPO/GRPO terminal reward, inspectable margins | Overoptimization, verbosity bias, Goodhart drift |
| Direct Preference Optimization (DPO) | Implicit policy/reference log-ratio | Trains a policy from labeled pairs without a separate scorer | Basic recipe depends on pair coverage; reference and data choices matter |
| Process Reward Model (PRM) | Step-level scalars | Dense credit assignment for multi-step reasoning traces | Expensive step-level labeling, still gameable[13] |
| Generative RM / LLM-as-judge | Verdict text or probabilities of verdict tokens, optionally reasoning | Flexible rubrics and synthetic label generation | Latency, position bias, shared factual errors; rationale need not be faithful[14] |
| Verifiable reward / RLVR | Execution or constraint-check result, often binary | Direct checks for specified properties | Incomplete tests, parsing errors, and gaps between checked properties and the real goal[5] |
- Process Reward Models (PRMs): Step labels identify where a reasoning trace goes wrong. Lightman et al. found process supervision improved best-of-N selection on MATH in their experiments. They evaluated a fixed generator and did not train it with RL; the result does not establish a universal advantage for all code or reasoning tasks.[13]
- RL with Verifiable Rewards (RLVR): Tulu 3 replaced a learned reward signal with answer or constraint verification for selected tasks.[5] This removes that learned scorer's errors, but reward hacking can still target an incomplete verifier. A function hard-coded to pass the visible tests may fail hidden inputs; an answer parser can accept a correct final number despite a false explanation.
Put the pieces together with TRL
The CPU head experiment isolates a mechanism. For a transformer run, TRL 1.13.0 assembles the sequence-classification model, candidate batching, and stable pair loss. This fragment assumes train_pairs and eval_pairs are Hugging Face Dataset objects containing validated pairs. They must already implement your intended split; the trainer does not group prompts for you.
1from trl import RewardConfig, RewardTrainer
2
3# Run with trl==1.13.0 and compatible dependencies.
4# train_pairs/eval_pairs are prepared Dataset objects, not defined here.
5trainer = RewardTrainer(
6 model="Qwen/Qwen3-0.6B",
7 args=RewardConfig(
8 output_dir="reward-head",
9 max_length=2048, # Example limit; audit filtering against your data.
10 center_rewards_coefficient=0.01, # Starting experiment, not a guarantee.
11 ),
12 train_dataset=train_pairs,
13 eval_dataset=eval_pairs,
14)
15trainer.train()
16metrics = trainer.evaluate()
17trainer.save_model("reward-head/final")
18trainer.processing_class.save_pretrained("reward-head/final")The trainer creates a one-label sequence-classification head when given this model ID. Set epochs, precision, batch size, and learning rate for your hardware and data rather than assuming the defaults are optimal. With LoRA, also save the trained score head and record the exact base revision. Keep the tokenizer and template with the scorer, and run held-out pairs after reloading the artifact.[6] This fragment downloads a model and trains it; it was source-checked, not executed as part of the CPU examples.
Before a small policy trial, inspect held-out and fresh-output results, calibration, annotation agreement, and shortcut slices. Define stop conditions using independent evaluation and a bounded optimization budget. A single accuracy threshold cannot certify the scorer.
Production diagnostic reference
When auditing training runs, check these symptoms and failure modes:
1. Training loss drops, held-out accuracy stalls
- Possible causes: Overfitting, noisy labels, easy-pair margin growth, or distribution differences. Stalled accuracy alone does not identify leakage.
- Diagnostic check: Audit overlap and label consistency, inspect held-out loss and calibration, and test controlled counterexamples.
2. Model consistently favors longer, wordier answers
- Possible causes: Length confounding, rubric preferences for detail, or presentation features correlated with length.
- Diagnostic check: Test length-matched and length-reversed pairs while preserving answer quality. Add counterexamples and re-evaluate; subtracting a length margin alone does not remove the shortcut.
3. Proxy reward climbs during PPO while human satisfaction collapses
- Possible cause: Proxy exploitation or changed evaluation/data distributions. Confirm the human metric and inspect the failing outputs.
- Diagnostic check: Stop the trial, preserve artifacts, and evaluate independently. A smaller budget, KL adjustment, improved labels, or a diverse ensemble may help; none repairs every evaluator error.
4. Raw reward values drift wildly across training steps
- Possible causes: Free score offsets, increasing margins, numerical errors, or inconsistent templates/pooling.
- Diagnostic check: Track pair midpoints and differences separately. Centering targets the offset; held-out calibration and regularization address different problems.