Policy gradients
A policy gradient is the gradient of expected reward with respect to the parameters of a stochastic policy. REINFORCE estimates it from sampled rollouts by weighting the gradient of each sampled action's log-probability by the reward that followed, so actions that led to better outcomes become more likely.
PPO and GRPO, the most common RL algorithms for LLMs, build on this estimator. Their changes affect how the weight is computed or how far one update may move the policy; the gradient itself stays recognizable.
The objective
A policy with parameters produces rollouts by interacting with an environment. Each rollout receives a reward . The objective is expected reward:
The difficulty is that appears in the sampling distribution, and the reward is typically a black box: a test suite, a verifier, a judge. There is no way to backpropagate through pytest.
The log-derivative trick
The identity moves the gradient inside the expectation:
The right-hand side is an expectation under the current policy, so it can be estimated by sampling. The reward only needs to be evaluated, never differentiated. This is also called the score-function estimator, since is the score.
The environment drops out
For an agent, a rollout interleaves the policy's actions with the environment's observations. Its probability factors as
where is the task, the policy's action at turn , the history before it, and whatever the environment does: run a command, return a search result, respond as a user. Taking the log turns the product into a sum, and only the policy terms depend on :
The environment's dynamics vanish from the gradient. The environment can be a sandboxed shell, a web browser, or another model, and it never needs to be modeled or differentiated. Its outputs still appear in the context that later actions condition on, but they contribute no gradient terms of their own. The policy-gradient loss implements this with a loss mask.
For a language model, each action is itself a sequence of tokens, so the sum runs over every sampled token in the rollout, each conditioned on its full context , observations included.
REINFORCE
Averaging over sampled rollouts gives REINFORCE (Williams, 1992):
This estimator is unbiased and very noisy. Two refinements are nearly always applied. Subtracting a baseline that does not depend on the sampled action leaves the expectation unchanged and, with a well-chosen baseline, reduces variance. This replaces with an advantage . And since an action cannot affect rewards that arrived before it, each token can be weighted by the reward that follows it (the reward-to-go). With the single terminal reward typical of LLM tasks, the reward-to-go at every token is , so this second refinement changes nothing. If the baseline is also one number per rollout, as in GRPO, every token of rollout shares the advantage . A learned critic that estimates value at each position would instead give each token its own advantage.
A policy-gradient update
What one update does to the probabilities
Take a single decision with three candidate tokens and softmax probabilities . The policy samples token 2 and the rollout receives advantage . For a softmax over logits , the gradient of the log-probability of the sampled token is
Gradient ascent with step size moves the logits by : the sampled token's logit rises and the others fall in proportion to their current probability. With every sign flips. With nothing moves, which is why uniform groups contribute no signal. Tokens that were already near probability 1 have and barely change, so most of the update lands on the uncertain decisions in a rollout.
In a full network the update goes through shared weights, so raising one sampled sequence's probability changes many other contexts too. A batch is a sum of such pushes, which can generalize, interfere, or concentrate the policy onto features shared across tasks.
In code
Automatic differentiation computes the estimator from a surrogate loss whose gradient equals , with advantages in place of rewards:
# logprobs: log π_θ(y_t | x, y_<t) for every token, shape [batch, seq]
# mask: 1 on tokens the policy sampled, 0 on prompt and observations
# advantages: one value per rollout, broadcast over its tokens
N = logprobs.shape[0]
loss = -(advantages[:, None] * logprobs * mask).sum() / N
loss.backward()This sums each rollout's token contributions and averages over rollouts, as in . Many implementations instead divide by the number of sampled tokens. That token-mean normalization is a different estimator, and the choice of denominator changes what is optimized, as covered under length bias. The value of the loss is not the objective and does not track learning progress; only its gradient matters.
The on-policy assumption
The derivation samples from , the same parameters being differentiated. Once the policy takes a step, the old rollouts come from a slightly different distribution. Reusing them, running several optimizer steps per batch, or generating asynchronously all break the assumption. On-policy and off-policy RL describes the distinction, and importance sampling and clipping are the corrections built on top of the plain gradient.