Entropy collapse and exploration
Entropy collapse is the rapid loss of diversity in a policy's outputs during reinforcement learning: token distributions become sharply peaked and rollouts for the same task start to look alike. Group-based RL learns by comparing different attempts at one task, so once the attempts stop differing, the training signal fades with them.
Policy entropy
The entropy of the policy at one position is
summed over the vocabulary. Training logs usually report its mean over sampled positions. A position with next-token probabilities has entropy 1.03 nats. After the policy commits to its favorite, , it has 0.15 nats.
Some decline is expected, because a policy that learns a task becomes more confident about it. Collapse is an early or abrupt drop that arrives with symptoms: repeated answers across a group, shorter or templated reasoning, tool calls that disappear, and a stalling reward curve.
Why it feeds on itself
A policy-gradient step raises the probability of actions with positive advantage and lowers those with negative advantage. Whether entropy falls depends on which actions win. If the actions the policy already favors are the ones earning positive advantage, probability piles further onto them, the alternatives are sampled less often, and they get fewer chances to prove themselves. If a low-probability action wins, the step spreads probability out and entropy rises.
Cui et al. turn this into a formula for a softmax policy with separate logits at each context. To first order, one policy-gradient step with step size changes the entropy at a context by
The covariance is positive when confident actions receive the advantage-weighted push, and then entropy falls. Take , which has entropy 0.90 nats, and one step with that moves each logit by :
| Which action won | Covariance | Entropy after the step | |
|---|---|---|---|
| The favorite | 0.85 nats | ||
| The rarest | 0.93 nats |
Once training works, the favorite is usually the one that wins, so the covariance stays positive step after step. Cui et al. also measured what this costs. Across their runs, downstream performance rose as entropy fell along the fitted curve . At the curve reaches its ceiling , so a policy that spends its entropy early stops improving early.
Group-relative methods like GRPO feel this directly. If every rollout in a group follows the same path, the rewards agree, the advantages are zero, and the group contributes nothing. It is the all-fail group of cold start reached from the other side.
Controls
The options differ in what they act on.
- Clip-higher. A symmetric ratio clip stops a rare token's gradient after a tiny absolute gain. Raising the upper bound of the clip gives rare tokens room to grow.
- Forking tokens. A minority of positions carry most of the entropy: words such as However or Wait where the reasoning can branch. Wang et al. restricted the gradient to the 20% of tokens with the highest entropy and matched full-gradient training on Qwen3-8B, gained 11.04 points on AIME'25 on Qwen3-32B, and lost substantial performance when training only on the lowest-entropy tokens. CISPO was motivated by the same tokens, which PPO-style clipping had been silencing, and a probability-space mask keeps rare tokens that a ratio bound would drop.
- Covariance-targeted regularization. Cui et al. act on the formula above. Clip-Cov removes the gradient from a small random subset of tokens with high covariance between log-probability and advantage, and KL-Cov puts a KL penalty toward the sampling policy on the highest-covariance tokens.
- Negative-sample reinforcement. Zhu et al. split the policy gradient into reinforcing correct rollouts and penalizing incorrect ones. Training only on the incorrect ones often matched or beat PPO and GRPO across pass@k, because pushing down a wrong answer returns its probability to alternatives the model already considered. Training only on the correct ones improved pass@1 and hurt large- pass@k by narrowing the output.
- pass@k as the objective. Rewarding a group when any of attempts succeeds removes the penalty on a diverse attempt that fails alongside a successful one. Chen et al. derive the advantages for this objective in closed form and report that it improves exploration, with exploration and exploitation reinforcing each other.
- An entropy bonus. Adding to the loss, with coefficient , pushes entropy up everywhere. It raises entropy on tokens that should be deterministic, such as code syntax or JSON structure, as much as on decision points.
- Data. Tasks the policy solves sometimes, and not always or never, keep groups mixed. Difficulty filtering is often the most effective control.
Raising the sampling temperature above 1 is not on the list. It produces more varied rollouts but leaves the model as peaked as before, and it changes the behavior policy , so the trainer learns from a distribution that differs from the one it is shaping and must importance-weight the difference (why ). A reference KL also holds entropy up, at the price of limiting how far the policy can improve.
Diagnosing collapse
Mean token entropy alone can mislead. It is dominated by the many tokens that should be nearly deterministic, and a model can keep high entropy on filler while collapsing on the few decisions that matter. Read it alongside:
- the fraction of groups whose rewards are all equal
- pass@k at several on a fixed evaluation set, since pass@1 can rise while large- pass@k falls
- response length and its variance
- the distinct final answers or distinct tool sequences per group
- entropy per environment in a multi-environment run, since one environment can collapse under a stable overall mean
- a sample of traces read directly
Tool-use collapse
Tool-use collapse is the same loop applied to actions: the policy stops calling tools that would help it.
Rollout-level credit often starts it. Early tool calls are frequently malformed or appear in failed rollouts, so they inherit negative advantage. If the model occasionally succeeds without the tool, avoiding calls becomes the locally safer strategy. Fewer calls produce fewer successful examples of tool use, which reinforces the decline.
Average reward can hide this for many steps. Track the call rate, calls per rollout, tool errors, and whether later actions use what the tool returned. Reading traces separates collapse from a policy that learned a shorter procedure.
The fix targets the cause: simplify the tool schema, make malformed calls recoverable instead of fatal, warm-start from successful tool traces, or give local credit to an intermediate result that can be verified. A small per-call cost may remove tool spam, while a large one can cause the collapse it was meant to prevent.