Rewards
A reward is a number that scores how well a rollout accomplished its task. In agent RL it is usually the main lever: the algorithm, the batch size and the learning rate determine how fast the policy moves, but the reward determines where it goes, and training follows the reward wherever it departs from the prompt. A coding task whose reward checks only that pytest exits with status zero is, to the optimizer, a task about making pytest exit with status zero, and deleting the failing test satisfies it.
Rewards come from two sources. A program can check the outcome by running tests, comparing a parsed answer or inspecting the final state of a database, which is the subject of verifiable rewards. When no program can decide the question, a learned model or prompted judge supplies the score. Many environments use both, and when training finds a path the designers did not intend, the result is reward hacking.
Outcome rewards
A rollout is a sequence of turns , each of which can carry a reward . RL maximizes the return, the sum of those rewards:
An outcome reward is nonzero only at the end: , and scores the finished result: whether the tests passed, the answer is correct or the database is in the required state. Most agent RL uses outcome rewards, because a correct final state is easier to define than a correct intermediate step. Rewards on intermediate steps are process rewards, which trade denser signal for proxies the policy can satisfy without making progress.
Worked example
A task often has several things worth checking. A math task might check the answer and a required output format. A coding task might check the target test, the rest of the suite, and whether the agent stayed within the files it was allowed to change. The usual combination is a weighted sum of component scores with weights :
Take a math task with a correctness score and a format score, each in , weighted and :
| Rollout | Correct | Formatted | Reward |
|---|---|---|---|
| A | 1 | 0 | 1.0 |
| B | 0 | 1 | 0.2 |
| C | 1 | 1 | 1.2 |
| D | 0 | 0 | 0.0 |
The ordering is C, A, B, D. Correctness dominates, and format breaks ties among correct answers and among wrong ones. Raise the format weight to and A and B tie, so a wrong answer in the right format earns as much as a right answer in the wrong one. Above the policy gains more from format than from correctness. A shaping term works as guidance only while its weight stays well below the gap between success and failure on the main outcome.
The components are worth recording separately even though the trainer consumes one scalar. A rising total can come from the main objective or from an auxiliary term, and only the components distinguish the two. A metric is recorded on every rollout without contributing to the reward (reply length, turn count, whether a protected file was touched), and it can be promoted to a reward later by giving it a weight.
Tool-call bonuses and penalties
A bonus per tool call teaches busywork: more searches, more commands and longer traces with no better outcome. A penalty per call expresses a cost the deployment pays, but it competes with the outcome. With success worth and a penalty of per call, a successful rollout that needs 25 calls scores , below a rollout that answers immediately and fails with . The policy learns to stop calling tools it needs, which feeds the tool-use collapse that rollout-level credit already encourages.
A cost stays a tie-breaker when it is small relative to the success gap and bounded. At per call, capped at , the same success scores and still beats every failure. Charging the cost only on successful rollouts, or logging call count as a metric instead of rewarding it, avoids the trade entirely. With a group-mean baseline, adding the same constant to every reward in a group leaves the advantages unchanged, so only differences between rollouts matter.
Partial credit
A binary reward gives a group no signal when every rollout fails, which is the uniform-group problem. Partial credit scores how close a rollout came. Exact match on a string-reversal task gives a weak model zeros almost everywhere. Scored by longest-common-subsequence ratio, a 20-character reversal with one swapped pair of letters earns , so rollouts in a group differ and the policy has something to climb. The cost is that partial credit pays for near misses. It fits tasks where a closer output is a better output, such as reversal, and misleads on tasks where an almost-correct answer is as useless as a wrong one, such as a numeric result or a patch that fails one test.
Rewarding abstention and honesty
A binary accuracy reward pays nothing for "I don't know" and nothing for a wrong answer, so any guess weakly dominates abstaining. Kalai et al. argue that this scoring rule, used across training and benchmarks, is a main cause of hallucination, and they prove that under binary grading abstaining is never optimal. Their fix is a confidence threshold : a correct answer earns , abstaining , and a wrong answer . With the penalty is . A model sure of its answer expects from guessing under binary scoring, and under the threshold, so abstaining becomes the reward-maximizing choice.
The agent version is an explicit way to report that a task cannot be completed. Without one, a coding agent facing a contradictory task can only fail or fake success, and faking scores higher whenever the verifier can be fooled. A separate honesty channel is another option: Joglekar et al. trained a model to write a confession after its main answer, rewarded only on whether the confession was honest, and found that when the model misbehaved in its answer, it confessed 74% of the time on average across their evaluations.
Rewards without labels
Some methods derive the reward from the policy itself. TTRL samples many answers to an unlabeled question, takes the majority vote as the pseudo-label, and rewards agreement with it; it raised Qwen2.5-Math-7B's AIME 2024 pass@1 by about 211% (relative) without ground-truth answers. Intuitor rewards self-certainty, the average KL divergence between a uniform distribution and the policy's next-token distribution, and roughly matched GRPO with answer rewards on math benchmarks. These rewards sharpen what the model already prefers, and Shao et al. found that even random rewards raised Qwen2.5-Math-7B on MATH-500 by 21.4 points while mostly failing on Llama3 and OLMo2. A gain from a label-free reward is convincing only when it holds on a second model family and an uncontaminated benchmark, the issues taken up in RL scaling.