Post-training methods
Post-training is everything done to a pretrained language model to shape its behavior: supervised fine-tuning (SFT), rejection sampling fine-tuning, online reinforcement learning, distillation, and preference optimization. An agent RL run is one stage in such a pipeline, and the stages before and after it determine what the policy can already do when RL starts and what survives after it ends.
Two questions separate the methods. The first is who wrote the tokens being trained on: the student, meaning the model being updated, or some other model or person. The second is what signal moves the weights at those tokens: the tokens themselves as targets, a filter, a reward, or a teacher's probabilities.
Two kinds of teacher
"Teacher" covers two different roles. A teacher can write the tokens: a person or a stronger model produces demonstrations, and the student is trained to reproduce them. SFT and hard distillation work this way, and the tokens are off-policy because the student did not sample them. A teacher can instead score the student's tokens: the student samples a rollout, and the teacher assigns a log-probability to every token the student chose. That is on-policy distillation.
| method | who wrote the tokens | signal at each token | pushes bad tokens down |
|---|---|---|---|
| SFT, hard distillation | a teacher or a dataset | the token itself as target | no |
| Rejection sampling fine-tuning | an earlier student | verifier accept or reject | no |
| Online RL | the current student | reward, turned into an advantage | yes |
| On-policy distillation | the current student | teacher log-probability | yes, per token |
| DPO and relatives | a fixed set of response pairs | preference comparison | yes, relative to the chosen response |
Rows where the current student writes the tokens train on the states the student reaches. Rows where someone else writes them train on the states that writer reaches, and the student has to generalize from there to its own.
Supervised fine-tuning on teacher tokens
In SFT a dataset supplies both the contexts and the target tokens. The loss is the negative log-likelihood of each target token given what came before:
Here is the prompt, is the demonstrated response, and the sum runs over the response tokens. Every token is pushed up, and nothing in the data says which alternatives were worse. Hard distillation is the same loss on responses a teacher model generated, so SFT is best read as a teacher for off-policy tokens, whether the teacher is a person or a model.
SFT is the standard way to install formats and protocols: a chat template, a tool-call syntax, a habit of ending with a final answer. For agents it usually comes before RL, because outcome rewards carry no signal while the model rarely emits a parseable tool call (cold start).
Its limitation is coverage. The demonstrations only show states the teacher reached. Once the student makes a mistake the teacher would not have made, it is in a context the data never covered, and in multi-step agent tasks a different file read or a failed command changes every later turn.
Rejection sampling fine-tuning
Rejection sampling fine-tuning lets the student write its own demonstrations. Sample several responses per prompt, keep the ones a verifier accepts, and run SFT on the survivors. Repeating the loop with the improved model gives expert iteration. STaR applied it to rationales, and Yuan et al. combined accepted samples from several models and raised LLaMA-7B to 49.3% on GSM8K, against 35.9% for SFT.
Consider eight samples for one prompt, two of which pass. Rejection sampling fine-tuning trains on those two with weight 1 and ignores the other six, which is a policy gradient with advantage equal to the raw reward. A group-relative method such as GRPO subtracts the group mean of 0.25, so the two passing samples get and the six failures get . The failures now carry signal: whatever they did is made less likely. Both methods use the same evidence, a verifier's verdict on the student's own outputs. Rejection sampling fine-tuning discards the negative half of it and usually trains in offline rounds, so its samples come from an older policy than the one being updated.
Online RL on the student's rollouts
Online RL samples fresh rollouts from the policy being trained, scores them, and updates toward the better ones. The score can come from a verifier or from a reward model trained on human comparisons, as in the classic RLHF pipeline of Ouyang et al.. The reward describes the behavior the model produces now, including strategies no demonstrator showed, and the policy visits the states its own choices create.
Training on its own samples also limits what RL overwrites. RL's Razor (Shenfeld et al.) found that forgetting tracks the KL divergence from the base policy on the new task, and that SFT on data generated by an RL-trained model matched RL's accuracy–forgetting trade-off, so the distribution trained toward, not the algorithm, sets how much is forgotten.
The costs are generation throughput and a sparse signal. A single scalar per rollout has to be spread over thousands of tokens (credit assignment), and a flawed reward is optimized as faithfully as a correct one (reward hacking).
Scoring the student's tokens with a teacher
On-policy distillation keeps RL's state distribution and replaces the sparse reward with a dense one. The student samples the rollout, a teacher computes its log-probability for each sampled token, and the gap between teacher and student log-probabilities becomes a per-token advantage. The teacher can be a larger model, a specialist, or the student itself shown privileged context such as a worked solution (self-distillation). It needs no reward function, and it teaches only behavior the teacher already assigns high probability.
Preference optimization
Preference methods learn from comparisons between two fixed responses. DPO fits the log-ratio between the policy and a frozen reference directly to preference pairs, with no reward model and no sampling during training. Identity Preference Optimization (IPO) replaces DPO's logistic loss with a squared one, and KTO uses unpaired good and bad labels. Relative to online RL, these are offline: the responses came from an earlier model, so the policy never learns what follows its own new behavior, and iterative variants that regenerate and relabel responses each round move back toward online RL.
Composing stages
Production pipelines chain these methods, and each stage prepares the next.
- Mid-training continues pretraining on curated data before any post-training. OctoThinker (Wang et al.) mid-trained Llama-3.2-3B on 200B tokens of math-heavy data and then 20B tokens of chain-of-thought data, after which RL brought it on par with Qwen2.5-3B, from a family that responds well to RL.
- SFT installs the output format, tool protocol and a starting level of competence, so that RL groups contain some successes.
- RL improves outcomes against rewards on the student's own rollouts.
- Consolidation and distillation move what specialists or a flagship learned into the model that ships. DeepSeek-V4 merges separately RL-trained domain experts into one student by on-policy distillation, and Qwen3 builds its small models from the flagship by off-policy distillation followed by on-policy distillation.