Policies
A policy is the rule an agent uses to choose actions given what it has observed. For an LLM agent, the policy is the model's probability distribution over the next token given its context, written , where are the weights that RL updates.
Every quantity an RL update uses, from the policy gradient to the importance ratio, is computed from this token distribution. Which distribution that is depends on the weights and also on the sampling settings that generated the data.
Tokens and actions
The environment never sees individual tokens. The harness lets the model generate until it returns a response or a tool call, parses the sampled tokens, and hands the resulting action to the environment. Model and harness together induce a distribution over commands, messages, patches and other agent actions.
A tool call as sampled tokens
Suppose the model emits a command as tokens after context . The probability of that sequence is the product of the per-token probabilities, and training code computes its logarithm as a sum:
The environment receives a parsed action, while the trainer evaluates the probability of the sampled tokens that encoded it. The two descriptions diverge in three ways.
- Many-to-one parsing.
{"cmd": "ls"}and{ "cmd":"ls" }are different token sequences that parse to the same action. The probability of the action is the sum over all its encodings, which no trainer computes, so training uses the probability of the one sequence that was sampled. - Rejection and repair. A parser may reject malformed output or quietly fix it, so the action that ran can differ from what the tokens said (tool calls).
- Behavior outside the weights. The system prompt, chat template, tool schema and stopping rules change the distribution of agent actions with the weights held fixed.
Sampling settings and the behavior policy
The weights and context determine the logits. Temperature, top-p, top-k and any other decoding rule then turn the logits into the distribution that generates data. That distribution is the behavior policy . The target policy is the one the loss optimizes. When the two match, training is on-policy, and each token's ratio equals 1 before the first gradient step (on-policy and off-policy RL).
Take three candidate tokens with logits . At temperature , the probabilities are .
| Temperature | |||
|---|---|---|---|
| 1.0 | 0.665 | 0.245 | 0.090 |
| 0.5 | 0.867 | 0.117 | 0.016 |
With identical weights, the third token is almost six times less likely at . Two inference workers serving the same checkpoint under different sampling settings generate data from different policies.
Truncated sampling changes the support as well as the shape. Top-p at zeroes every token outside the smallest set holding 90% of the mass and renormalizes the rest, so a sampled token has a higher probability under than under the raw model, and tokens outside the set are never sampled. A change of temperature can be corrected by importance weighting, because both distributions give every token some probability. Truncation cannot be corrected that way, since the samples say nothing about the removed tokens. The workable option is to replay the recorded truncation set in the trainer, which makes the target a policy restricted to that support.
Either way, the per-token log-probabilities stored for training should come from the distribution that sampled each token, temperature and truncation included. Any other difference between the sampling engine and the trainer's forward pass also makes differ from (trainer–inference mismatch).
Why sample at temperature 1.0?
Group-based methods such as GRPO learn from differences between rollouts of the same task. With greedy decoding every rollout would be identical, every advantage zero, and the update empty. The variance across rollouts is the learning signal, and temperature scales how much of it the policy produces.
Suppose the third token in the table is the start of a rare but better strategy, such as reading a file before editing it. In a group of 16 rollouts, the chance that at least one rollout samples it is at and at . Lowering the temperature makes a group more likely to agree, and a group that agrees on its reward carries no signal.
Lowering the temperature also changes the behavior policy. If the trainer computes at while rollouts were sampled at , the ratio compares two different distributions and the gap looks like off-policy drift. If the trainer applies the same temperature, the loss optimizes the tempered policy, and deployment at another temperature runs a policy that was never trained. Sampling at keeps the behavior policy equal to the model's own distribution. Evaluation may use other settings, which should be reported with the results (evaluation).
The policy's randomness shrinks as training sharpens it, whatever the sampling temperature. When entropy collapses, rollouts of the same task become near-copies and learning stalls (entropy collapse). Prompt and harness changes alter the behavior policy without touching the weights, which is why they can be optimized alongside RL (harnesses).