On-policy distillation
On-policy distillation trains a student model on rollouts the student samples itself, with a teacher model scoring every token the student chose. For agent RL it supplies a dense signal on the states the student reaches, where an outcome reward gives one number per rollout.
The common recipe minimizes the per-token reverse KL from the student's next-token distribution to the teacher's. Other variants match more of the teacher's distribution or use a different divergence.
Why teacher-written rollouts fall short
The simplest way to distill is to have the teacher write rollouts and run SFT on them. The contexts then come from the teacher, so the student learns what to do in states the teacher reaches and never sees what to do after its own mistakes. GKD (Agarwal et al.) and MiniLLM (Gu et al.) addressed this mismatch by training on student-generated sequences.
On-policy distillation keeps the student's rollout and asks the teacher only to score it. If the student reads the wrong file at turn three, the teacher's log-probabilities from turn four onward describe how a stronger model would continue from that wrong turn. That recovery behavior never appears in a demonstration the teacher wrote itself.
The per-token reverse-KL signal
For a sampled token in context , with student and teacher , define the per-token distillation signal
Because was sampled from the student, the expectation of at that position is the reverse KL . Minimizing it with a policy gradient treats as a per-token advantage, held fixed during the update: tokens the teacher likes more than the student are pushed up, and tokens the student overrates are pushed down. The Thinking Machines write-up that popularized the recipe computes this signal from teacher log-probabilities on student samples.
The update treats each recorded context as fixed. It lowers the KL at positions the student already visited, but ignores how changing the student changes which contexts it reaches later, so it is not the full gradient of the sequence-level KL. The per-token KL penalty has the same partial gradient.
Only the teacher's log-probability of the sampled token is needed. The teacher runs a prefill over the student's tokens with no generation, which costs far less than sampling and needs no access to the teacher's full vocabulary distribution.
Worked example
Take three tokens from one student rollout.
| token | effect | |||
|---|---|---|---|---|
| 1 | 0.60 | 0.62 | +0.03 | almost no change |
| 2 | 0.40 | 0.05 | −2.08 | strongly pushed down |
| 3 | 0.10 | 0.50 | +1.61 | pushed up |
An outcome reward would give all three tokens the same advantage. Here token 2 is singled out as the mistake even if the rollout ended correctly, and token 3 is reinforced even if it ended in failure.
Reverse KL versus forward KL
Classic distillation minimizes the forward direction, , averaged over teacher contexts, and SFT on teacher samples is a single-sample estimate of it. Forward KL is mode-covering: the student is penalized for giving low probability to anything the teacher might say, so a small student spreads mass across all of the teacher's options.
Reverse KL is mode-seeking. The student is penalized for saying things the teacher would not, and not for ignoring some of the teacher's options. A student with less capacity than its teacher can commit to one of the teacher's behaviors instead of averaging several. The cost is lower diversity across rollouts, which reduces the variance later RL learns from (entropy collapse).
Combining with RL
Distillation trains agreement with a teacher, not task success, so its signal is only as good as the teacher's behavior on the task. It does not cap the student at the teacher's score: Born-Again Networks trained image classifiers to match a teacher of identical architecture and found students that outperformed it. RL with a verifier adds a signal for successful behavior the teacher rarely produces. Where a strong teacher exists, distillation is also much cheaper per unit of improvement: starting Qwen3-8B from the same off-policy-distilled checkpoint, Qwen3 reports 74.4 on AIME'24 (accuracy averaged over 64 samples) from on-policy distillation after 1,800 GPU hours, against 67.6 from RL after 17,920.
The two signals can run in sequence or jointly, with both losses optimized in the same run. Sequential Beats Joint (Li et al.) found that on-policy distillation followed by RL outperformed pure distillation, pure RL and every joint baseline it tested on logic and math reasoning. Distillation widens the student's coverage of solutions the teacher supports, RL sharpens behavior within that coverage, and optimizing both at once lets the signals interfere. The authors used the distillation validation score to decide when to switch, and found distillation a better cold start for RL than SFT.
Several teachers can feed one student. The per-token signal is unchanged, and only the model computing varies with the task. DeepSeek-V4 uses this to consolidate separately trained domain experts, such as mathematics, coding and agent specialists, into one model with a reverse-KL objective against all of them. The student learns each domain on its own state distribution, and no single run has to balance several reward scales.
Failure modes
- Tokenizer mismatch. The teacher has to score the student's exact token sequence. A teacher with a different tokenizer needs a bridge between vocabularies, which adds approximation.
- Truncated sampling. Under top-p or top-k sampling the tokens come from a truncated behavior policy, not from . The teacher scores them with full-vocabulary log-probabilities while the student's are renormalized over the truncated set, which biases .
- Unfamiliar states. A student far from the teacher reaches contexts where the teacher's predictions are poorly calibrated, and the teacher gives confident advice about situations it has little experience with.
- Masking. Only tokens the student sampled are trained. Tool outputs and other observations remain context under the same loss masks RL uses.
- Teacher errors. A confident teacher can be wrong. The student has to be evaluated on the capability of interest, not on agreement with the teacher.