Reward models, judges, and rubrics
A reward model is a learned model that scores a response, used as the reward when no program can check the outcome. An LLM judge does the same job by prompting a capable model with the task, the response and grading instructions. Both extend RL to qualities such as helpfulness, faithfulness to sources or the quality of a written report, where verifiable rewards are unavailable, and both are predictions that a policy under optimization will probe for errors.
Preference reward models
Reinforcement learning from human feedback (RLHF) trains a reward model on human comparisons and then optimizes a policy against it. Christiano et al. introduced it for control tasks, and InstructGPT scaled it to instruction following.
The data are pairs: for a prompt , a labeler marks a preferred response over a rejected one . The reward model outputs a scalar, and the Bradley–Terry model turns the difference between two scalars into a preference probability:
where is the logistic function. Training minimizes .
If the model scores the chosen response and the rejected one , the difference is , the predicted preference probability is , and the loss is . The gradient pushes the two scores further apart. Only differences enter the loss, so the offset of a reward model's outputs is arbitrary and has to be fixed, for example by normalization, before the score is combined with other reward components.
Overoptimization
A reward model is accurate on the kind of responses it was trained to compare, and RL moves the policy away from that distribution. Gao et al. measured the effect with a large "gold" reward model standing in for humans: as a policy is optimized against a smaller proxy reward model, the proxy score keeps rising while the gold score rises, peaks and then falls. Larger proxy models reached higher peak gold scores and overoptimized less, with the fitted decline shrinking smoothly as proxy size grew. This is reward-model overoptimization, the form reward hacking takes with learned rewards.
The usual controls are a KL penalty that limits how far the policy moves from its starting point, refreshing the reward model with comparisons between responses from the current policy, and ensembles of reward models whose errors differ. Each slows the drift into regions where the proxy fails, and none removes it, so a learned reward needs a gold-standard check (human labels or a held-out verifier) at intervals during training.
Outcome and process reward models
An outcome reward model (ORM) scores a finished response, typically by predicting whether it is correct. Cobbe et al. trained such verifiers on GSM8K solutions to rerank samples. A process reward model (PRM) scores each intermediate step, the learned form of the process rewards used for dense credit. Lightman et al. collected 800,000 step-level human labels (PRM800K) and found that a PRM selected correct solutions on MATH more reliably than an ORM (78% against 72% of problems solved, best of 1,860 samples). A solution's PRM score is usually the product or minimum of its step scores.
Step labels are expensive, a step must be well defined, and a policy trained against step scores can learn steps that look good to the PRM without leading anywhere. For agents the natural step is a turn, and whether a turn was good often depends on what happened several turns later, which is why agent RL mostly scores outcomes.
LLM judges
An LLM judge replaces a trained scoring head with a prompt. It receives the task, the response, and usually a reference answer or grading criteria, and returns a verdict. Zheng et al. found that GPT-4 agreed with human preferences on open-ended chat over 80% of the time, the same rate at which humans agree with each other, and documented position bias, verbosity bias and self-preference. Wang et al. showed that swapping the order of two candidates can flip a judge's preference.
Judges are most stable with a short prompt aimed at one decision:
- Evidence. A reference answer, the relevant source documents, or the task's hidden requirements turn an opinion into a comparison.
- A small verdict space. Yes/no or a short ordered scale varies less between calls than a score out of ten.
- Recorded calls. Judge prompts, replies and costs belong in the trace, where judge errors can be found.
- Human agreement. Agreement with human labels on training rollouts, including high-reward ones from late in training, is the direct measure of whether the judge can be trusted.
Rubrics as rewards
A rubric breaks one judgment into criteria written for the specific prompt, each graded separately by a judge and combined with weights. Take three yes/no criteria with weights , and : the answer is correct, it cites sources, it uses the requested format. A response that is correct and cites sources but misses the format scores . The per-criterion verdicts show whether a rising score came from correctness or from citations, which a single holistic score hides.
Gunjal et al. trained with per-prompt rubrics on medicine and science questions and improved on a judge giving a single Likert score by up to 31% relative on HealthBench. Viswanathan et al. extracted a checklist of requirements from each instruction, graded items with a judge or a verification program, and raised Qwen2.5-7B-Instruct by 4 points on FollowBench and 6 on InFoBench. Huang et al. scaled the approach to more than 10,000 rubrics for open-ended tasks such as humanities writing. Each criterion is still a judge call, so each can be gamed on its own: "cites sources" rewards citations whether or not they support the claim, unless a criterion checks that too.
Generative reward models
A generative reward model writes a critique or reasoning before its score, instead of emitting a scalar from a classification head. DeepSeek-GRM generates its own principles and critique for each query, trained with online RL, and gets better by sampling several critiques and voting, so it can trade inference compute for accuracy. RM-R1 generates a rubric or reference solution first and then judges against it, and outperformed GPT-4o and a 70B scalar reward model by up to 4.9% on reward-model benchmarks. For RL, the cost is a long generation per score on every rollout, and the benefit is a reasoning trace that shows why a verdict was given.
Agent judges
Some outcomes can only be checked by doing work: running the code, querying the database, reading the files that changed. An agent judge is a judge run as its own rollout in a two-agent episode, like those in self-play: it gets tools and a runtime, receives the solver's trace and artifacts, investigates, and writes a verdict. A solver that says "all tests pass" can be checked by a judge that runs them. The costs are those of any rollout: more tokens and time, more variance, and a larger surface for manipulation, since the solver's output is untrusted input to the judge. Running the judge in a fresh runtime that receives only declared artifacts limits what the solver can plant.
A program is the preferred scorer wherever it can decide the outcome, because it is cheaper, deterministic and harder to manipulate. Many environments pair a verifier for correctness with a judge or rubric for properties such as clarity, and the more weight the learned score carries, the more of its high-scoring rollouts need human review.