PPO clipping and trust regions

PPO clipping is the mechanism in Proximal Policy Optimization that caps how much a sampled action can improve the objective: its contribution stops growing once its probability ratio rises above 1+ϵ1+\epsilon with a positive advantage, or falls below 1−ϵ1-\epsilon with a negative one. Gradients in the corrective direction are left in place. It is a cheap stand-in for a trust region: a limit on how far one update may move the policy away from the behavior policy that produced the data.

Why a trust region is needed

PPO was designed to take several epochs of minibatch updates on each batch of rollouts. After the first minibatch, the policy being trained no longer matches the one that sampled the data, so every later step is off-policy. The importance-weighted surrogate ρtAt\rho_t A_t, with ρt=πθ(yt∣x,y<t)/μ(yt∣x,y<t)\rho_t = \pi_\theta(y_t \mid x, y_{<t})/\mu(y_t \mid x, y_{<t}), adjusts action probabilities at the contexts that were recorded, which is a good approximation while the two policies stay close. It does not correct the distribution of whole rollouts. It also invites the optimizer to push ρt\rho_t as far as it likes on tokens with positive advantage, extrapolating from a single sample into regions the batch says little about.

Earlier work enforced a trust region with an explicit KL constraint, which requires second-order optimization. PPO replaced it with a clipped objective that plain SGD or Adam can optimize. The trust region anchors each update to the behavior policy that sampled the batch, which is a different anchor from the fixed reference model of a KL penalty.

The clipped surrogate

PPO's clipped surrogate for one token is

Ltclip(θ)=min⁡ ⁣(ρtAt, clip⁡(ρt, 1−ϵ, 1+ϵ) At),L^{\text{clip}}_t(\theta) = \min\!\Big(\rho_t A_t,\ \operatorname{clip}(\rho_t,\,1-\epsilon,\,1+\epsilon)\,A_t\Big),

maximized (or its negative minimized) over the batch. The original paper uses ϵ=0.2\epsilon = 0.2. Taking the minimum of the raw and clipped terms makes the objective a pessimistic bound, and its effect depends on the sign of the advantage:

  • At>0A_t > 0. Raising the token's probability helps until ρt=1+ϵ\rho_t = 1+\epsilon. Beyond that the objective is flat and the token's gradient is zero. Lowering the probability is never clipped, so a mistake in the harmful direction is always corrected.
  • At<0A_t < 0. Lowering the probability helps until ρt=1−ϵ\rho_t = 1-\epsilon, then the gradient is zero. Raising it is never clipped.

The unclipped side has no ceiling for negative advantages. A token with At<0A_t < 0 whose ratio has already grown to 50 still receives the full ρtAt\rho_t A_t gradient. INTELLECT-2 (§3.4) adds an upper bound δ>1+ϵ\delta > 1+\epsilon on the ratio for negative-advantage tokens, after tracing loss and gradient-norm spikes, which grew worse as its models got larger, to that unbounded side.

Clipped probability ratios

interactive
0.81.2
unclipped surrogatemin⁡(rtA^t,clip⁡(rt,0.8,1.2)A^t)=1.344\min\left(r_t \hat{A}_t, \operatorname{clip}(r_t, 0.8, 1.2) \hat{A}_t\right) = 1.344
Clipping caps improvement to this surrogate objective outside the interval. It does not impose a hard bound on the policy update.

A worked example

Take ϵ=0.2\epsilon = 0.2 and four tokens after a few optimizer steps on the same batch:

TokenAtA_tρt\rho_tρtAt\rho_t A_tClipped termObjectiveGradient flows?
1+11.101.101.101.10yes
2+11.501.501.201.20no, gain is capped
3+10.600.600.800.60yes, pushed back up
4−10.70−0.70−0.80−0.80no, penalty is capped

Token 2 has already gained as much as the clip allows, so the optimizer stops pushing it. Token 3 is a good token whose probability fell, and the unclipped branch keeps pulling it back. Token 4 is a bad token already pushed down past 0.80.8, and it stops receiving gradient.

What clipping does not do

Clipping removes the incentive to move a token further. It does not constrain where the token ends up. The parameters are shared across every token in the batch, so gradients from unclipped tokens can still push a clipped token's ratio to 2 or 0.3. How far the policy moves depends on the learning rate, the number of epochs and minibatches, the advantage scale, and the batch. A token's final ratio can therefore lie well outside the interval even though its own gradient stopped at the boundary.

Clipping also gives zero gradient to the tokens where the policy changed most in a favorable direction. For LLMs these are often rare tokens that the update is trying to make more likely, which connects clipping to entropy collapse.

Clip-higher and asymmetric intervals

DAPO decouples the two bounds into [1−ϵlow,1+ϵhigh][1-\epsilon_{\text{low}}, 1+\epsilon_{\text{high}}] with ϵlow=0.2\epsilon_{\text{low}} = 0.2 and ϵhigh=0.28\epsilon_{\text{high}} = 0.28, a change it calls clip-higher. The reasoning is that a symmetric upper bound is very tight in absolute terms for low-probability tokens. A token at probability 0.01 can only reach 0.012 before clipping, while one at 0.9 can reach 1.0. Raising the upper bound lets unlikely exploration tokens grow, which DAPO found counteracts entropy collapse.

Other methods keep a trust region but change what is bounded. CISPO clips the importance weight instead of the objective, so every token keeps a gradient. GSPO clips a length-normalized sequence ratio. Masks such as IcePop and the probability-difference mask drop out-of-range tokens outright.

Reading clip fraction

The clip fraction is the share of tokens whose ratio lies outside the interval, sometimes counted only where the clip zeroes the gradient. Alongside it, an approximate KL between the behavior and current policy on the batch shows how far the update moved.

The cause of a high clip fraction determines the fix:

  • High only on later epochs. The optimizer is outrunning the data. Lower the learning rate or take fewer epochs.
  • High on the very first step, before any update. The ratio should be 1. A large deviation means stale rollouts, trainer–inference mismatch, or a ratio computed for the wrong token context, for example after re-tokenization. More clipping cannot fix a ratio computed against the wrong prefix.
  • Rising steadily over a run. Policy lag or numerical mismatch is growing. Check staleness and inference settings.