Advantages and baselines
An advantage measures how much better a sampled action or rollout turned out than the policy should have expected in the same situation. It is the weight a policy-gradient update multiplies into each sampled token's log-probability, so its sign sets which agent behavior gets repeated and its magnitude sets how hard each rollout pushes.
The baseline is the expectation the reward is compared against.
Why rewards need a reference point
Suppose an agent solves 90% of the tasks in a batch and every success earns reward 1. With raw rewards as weights, every successful rollout is pushed up with the same strength, whether the task was trivial or the success was a lucky outlier, and every failed rollout receives zero weight.
Subtracting a baseline sets the reference point. A success on a task the policy almost always solves is expected, so its advantage is small. A success on a task it usually fails carries a large positive advantage, and a failure on an easy task a large negative one.
Baselines and the expected gradient
A baseline that depends on the state but not on the sampled action leaves the expected policy gradient unchanged:
Any such is allowed, and what it changes is the variance of the estimate. The conventional choice is the state value , the expected return from under the current policy. With that baseline the advantage estimates , how much taking action changes the expected outcome relative to the policy's average from . Centering each state's returns this way removes the variation that comes from some tasks being easier than others. is not the variance-minimizing baseline in general, since the optimum weights each return by the squared norm of its score function, but can be estimated from returns alone.
The derivation also sets a constraint. If the baseline uses the sampled action's own reward, the zero-mean argument breaks and the estimator becomes biased, which group baselines below have to handle.
Critics and GAE
A value model (or critic) is a learned network that predicts from the current context. PPO in its classic RLHF form trains one alongside the policy, initialized from a similar LLM and fit by regression to observed returns. With a critic, the per-step advantage can be built from temporal-difference errors:
Generalized advantage estimation interpolates between two estimators with . At the advantage is the one-step error , low variance but only as accurate as the critic. At the sum telescopes to the Monte Carlo return minus , unbiased but noisy. The PPO paper's reference settings use .
A critic costs a second model of similar size, a second optimization problem, and a new failure mode: a poorly fit critic produces confident, wrong advantages. In a long agent rollout the context holds tool outputs from a sandbox whose state the critic cannot inspect, which makes predicting success from a prefix hard. What per-turn values buy when they are accurate is the subject of credit assignment.
Undiscounted returns
Classic RL discounts future reward by because episodes can be long or unbounded. LLM and agent tasks are finite: a rollout ends when the model stops, the harness hits a turn or token limit, or the task terminates, and most tasks give one reward at the end. Discounting per token would erase that reward. A reward 10,000 tokens away, discounted with per token, receives weight , so LLM RL almost always sets .
With one terminal reward and , every sampled token in rollout shares the return . A rollout-level baseline gives all of them the same advantage , while a critic can give each position its own.
Baselines from a group of rollouts
Because the return is one number per rollout, the baseline only needs the policy's expected reward on that task. Sampling several rollouts of the same task estimates it without training anything.
Leave-one-out (RLOO). For rollouts with rewards , each rollout's baseline is the mean of the others:
The baseline excludes , so it does not depend on rollout 's actions and the estimator stays unbiased. Ahmadian et al. found that REINFORCE with this baseline outperforms PPO for RLHF at much lower cost.
Group mean. Subtracting the mean of all rewards, including , is the baseline GRPO uses.
Worked example. Four rollouts of one task earn rewards .
| Rollout | Reward | Group-mean advantage | RLOO advantage |
|---|---|---|---|
| 1 | 1 | ||
| 2 | 0 | ||
| 3 | 0 | ||
| 4 | 1 |
Every RLOO value is times the group-mean value. The ratio holds for any group:
Mean-centering is RLOO scaled by , so its expected gradient points in the same direction and is shorter by that factor. If all four rewards had been 1, both estimators would give zero everywhere, which is why difficulty filtering matters for group baselines.
Running baselines. A baseline can also be an exponential moving average of past rewards, per task or per agent role. It works with one rollout per prompt and suits multi-agent games, where a group mean over both sides of a zero-sum game is near zero whatever the policy does (self-play). The average trails a policy that is improving, so its advantages run slightly high during learning.
Baselines in agent RL
Group baselines dominate LLM RL with verifiable rewards. Generation is already the expensive part of a step, and rollouts per task add generation but no second model, while a critic adds a model that has to track the policy and still misjudges long tool-heavy prefixes. The price is resolution: a group baseline gives one advantage per rollout, and how that scalar is spread over turns is a separate choice.
A stronger baseline makes learning more selective but cannot repair the reward. If the scorer pays for a shortcut, the best shortcut in each group receives the largest advantage (reward hacking).