Data sampled by a recent, bounded-staleness version of the policy with recorded per-token probabilities, so a token-level importance correction can close the gap. On-policy and off-policy RL →
Admission control
Limiting how many rollouts are in flight based on the inference engine's cache usage and queue depth. Rollout generation →
Advantage
How much better a sampled rollout or action did than the policy's expected outcome, computed as reward minus a baseline. Advantages and baselines →
The decision-maker in RL; for an LLM, the model together with the harness that turns its tokens into actions. Agents and environments →
Agent judge
A judge run as its own rollout, with tools and a runtime, that verifies another agent's work before writing a verdict. Reward models, judges, and rubrics →
Anchor state
An environment state reached in several rollouts of a group, used to compare the actions taken from it. Credit assignment →
Asynchronous RL
Running rollout generation and training concurrently, so batches come from a slightly older policy; the lag is bounded and corrected by importance weighting. Asynchronous RL →
Baseline
A quantity subtracted from the reward to reduce policy-gradient variance; one that does not depend on the sampled action leaves the expected gradient unchanged. Advantages and baselines →
Batch invariance
The property that a kernel's output for one sequence does not depend on what else is in the batch. Trainer–inference mismatch →
Behavior policy
The distribution that sampled the data, written μ, including weights, temperature, truncation and inference numerics. Policies →
Bradley–Terry model
A preference model that turns the difference between two reward scores into the probability that one response is preferred, via a logistic function. Reward models, judges, and rubrics →
Branch
A root-to-leaf path through a trace's message graph; each trainable branch becomes one training sample. Rollouts and traces →
Cache salt
A per-request key that namespaces prefix-cache entries so rollouts from one policy version never reuse cache computed by another. KV cache and prefix reuse →
The share of tokens whose ratio falls outside the clipping interval, used to diagnose policy drift or mismatch. PPO clipping and trust regions →
Clip-higher
DAPO's asymmetric clipping interval with a larger upper bound, letting low-probability tokens grow. PPO clipping and trust regions →
Clipped surrogate objective
PPO's objective, which caps a sampled token's improvement once its ratio exceeds 1+ε with positive advantage or falls below 1−ε with negative advantage. PPO clipping and trust regions →
Cold start
The situation where the initial policy almost never succeeds, so groups fail uniformly and give no group-relative learning signal. Cold start and difficulty filtering →
Selecting or reweighting training tasks by the policy's estimated pass rate so that groups have mixed outcomes. Cold start and difficulty filtering →
Direct Preference Optimization (DPO)
An offline objective that fits a policy's implicit reward, relative to a frozen reference, directly to preference pairs. Post-training methods →
Dr. GRPO
A GRPO variant that removes std normalization and per-response length normalization to avoid difficulty and length biases. Group-relative advantages (GRPO) →
Dual-control environment
An environment in which both the agent and the simulated user hold tools that act on shared state. RL environments →
Dynamic sampling
Over-sampling prompts and discarding groups whose rewards are all identical, resampling until each batch is full of informative groups. Cold start and difficulty filtering →
ECHO
An objective that adds a weighted cross-entropy loss on environment observation tokens to the usual policy-gradient loss on actions. The policy-gradient loss →
Entropy collapse
The rapid loss of diversity in a policy's outputs during RL, which removes the variation group-relative learning needs. Entropy collapse and exploration →
Environment
Everything the agent acts on: the task, its external state, the processes that respond to actions, and the scorer. Agents and environments →
Episode
A container for rollouts that belong together, such as a solver and its judge or the agents in a multi-agent run. Rollouts and traces →
Expert iteration
Repeating rounds of sampling, filtering by a verifier or search, and fine-tuning on the successes. Post-training methods →
Forking tokens
The minority of high-entropy tokens where a model's reasoning can branch; restricting the policy gradient to them matches training on all tokens. Entropy collapse and exploration →
GAE
Generalized advantage estimation, which blends one-step temporal-difference errors and Monte Carlo returns with a parameter λ. Advantages and baselines →
An advantage computed by comparing a rollout's reward with the rewards of other rollouts sampled for the same task. Group-relative advantages (GRPO) →
GRPO
Group Relative Policy Optimization, which computes each rollout's advantage relative to the mean (and originally the std) of rewards in its group. Group-relative advantages (GRPO) →
Fine-tuning a student on rollouts a teacher wrote, which trains it on off-policy tokens. Post-training methods →
Harness
The program around a model that builds its requests, exposes tools, parses output into actions, handles errors, and decides when to stop. Harnesses →
IcePop
A trainer–inference calibration method that masks tokens whose training-to-inference probability ratio falls outside a two-sided band. Importance sampling →
Identity Preference Optimization (IPO)
An offline preference objective that replaces DPO's logistic loss with a squared loss to limit overfitting to deterministic preferences. Post-training methods →
Importance ratio
The probability of a sampled token under the current policy divided by its probability under the behavior policy. Importance sampling →
In-flight weight update
Swapping new weights into running inference requests so a single rollout spans several policy versions. Asynchronous RL →
Inoculation prompting
Framing an unwanted behavior as permitted in the training prompt so that learning it does not generalize to other settings. Reward hacking →
Interleaved thinking
Keeping a reasoning model's earlier thinking blocks in the context of later turns instead of dropping them at each new user message. Token-level rendering →
IPO loss (prime-rl)
prime-rl's default online RL loss, Importance Policy Optimization: an importance-weighted policy-gradient term with a symmetric probability-space mask. Importance sampling →
k3 estimator
For x sampled from q and r = p(x)/q(x), the nonnegative estimate (r − 1) − log r, whose expectation is KL(q ‖ p) when p and q have the same support. KL penalties and reference models →
KV cache
The per-layer store of attention keys and values that lets an inference engine extend a context without recomputing it. KV cache and prefix reuse →
Learnability
A proposer reward of 4p(1−p), where p is the solver success rate; it peaks when half the solvers succeed. Self-play and multi-agent RL →
Length bias
A distortion where per-sequence loss normalization penalizes long wrong answers less per token, pushing incorrect responses to grow longer. The policy-gradient loss →
A per-token flag that limits the policy-gradient loss to tokens the policy sampled, keeping prompts and observations as context only. The policy-gradient loss →
MaxRL
An advantage estimator that divides the group-centered reward by the group mean, targeting the likelihood of success rather than expected reward. Group-relative advantages (GRPO) →
Metric
A per-rollout measurement recorded alongside the reward but given no weight in it, used to detect problems the reward misses. Rewards →
Mid-training
Continued pretraining on curated data, such as math and chain-of-thought text, that prepares a base model for post-training and RL. Post-training methods →
Multi-adapter serving
Hosting many LoRA adapters on one base model and batching requests that use different adapters together. LoRA and parameter-efficient RL →
Multi-teacher consolidation
Merging separately trained specialist models into one student by on-policy distillation, each teacher scoring its own domain. On-policy distillation →
Negative-sample reinforcement
Training only by pushing down incorrect rollouts, which returns probability to alternatives and preserves pass@k at large k. Entropy collapse and exploration →
Off-policy data
Data sampled from a policy other than the one being updated, such as an older checkpoint or a different model. On-policy and off-policy RL →
On-policy distillation
Training a student on sequences it generates itself, with a teacher supervising its predictions at the contexts it visits. On-policy distillation →
OPSD
On-Policy Self-Distillation, where the teacher view of the same model sees privileged information such as a verified reasoning trace. Teachers and privileged context →
Outcome reward
A reward assigned only at the end of a rollout, based on the finished result. Rewards →
Partial credit
A graded reward that scores how close a rollout came to success, giving groups signal where a binary reward would be uniform. Rewards →
Partial rollout
A long generation capped at a per-step token budget and continued in a later step, so one rollout spans several policy versions. Rollout generation →
pass@k
The probability that at least one of k independent attempts at a task succeeds, estimated without bias from n ≥ k samples. Evaluating agents →
pass^k
The probability that all k independent attempts at a task succeed, a measure of an agent's consistency. Evaluating agents →
Policy
The model's probability distribution over the next token given its context, which RL updates. Policies →
Policy lag
The number of policy versions between the weights that sampled a rollout and the weights its batch updates. Asynchronous RL →
A Markov decision process in which the agent receives observations that reveal only part of the underlying state. Agents and environments →
PPO
Proximal Policy Optimization, which pairs a learned value baseline with a clipped importance-ratio objective to reuse a batch for several updates. RL algorithms for LLMs: PPO, GRPO, and variants →
Prefill–decode disaggregation
Running prompt processing and token generation on separate inference pools, transferring the KV cache between them. Inference-heavy workloads →
Prefix break
A point in a multi-turn rollout where the next prompt no longer extends the previously sampled tokens, which starts a new training sample. Token-level rendering →
Prefix caching
Reusing stored KV-cache blocks across requests that begin with the same exact token prefix. KV cache and prefix reuse →
Privileged context
Information such as a demonstration or verified solution given to the teacher but withheld from the student. Teachers and privileged context →
Probability-space mask
A trust region that drops tokens whose probability changed by more than a threshold in absolute terms, as in prime-rl's ipo loss. Importance sampling →
Process reward
A reward assigned to intermediate steps of a rollout, such as a verified lemma or a completed subgoal. Credit assignment →
Process reward model (PRM)
A reward model that scores each intermediate step of a solution, in contrast to an outcome reward model (ORM) that scores only the final result. Reward models, judges, and rubrics →
Prompt optimization
Improving an agent by searching over its prompt or harness with rollouts and reflection instead of gradient steps. Harnesses →
The basic policy-gradient estimator that weights the gradient of each sampled action's log-probability by the reward that followed. Policy gradients →
Rejection sampling fine-tuning
Sampling several responses, keeping those a verifier accepts, and fine-tuning on the survivors; a policy gradient whose advantage is the raw 0/1 reward. Post-training methods →
Renderer
The component that converts between chat messages and model token IDs, including parsing sampled output and building the next turn's prompt. Token-level rendering →
Reverse KL
The divergence KL(student ‖ teacher), estimated on student samples; it is mode-seeking. On-policy distillation →
Reward
A number that scores how well a rollout accomplished its task, and the working definition of success for RL. Rewards →
Reward hacking
Raising reward through behavior the designers did not intend, by exploiting a gap between the reward and the goal. Reward hacking →
The effect where optimizing against a learned reward keeps raising its score while true quality peaks and then falls. Reward models, judges, and rubrics →
RL environment
The executable definition of a task (tasks, tools, scoring, and the runtime they act on) that sets the starting state, actions, termination, and reward; the harness belongs to the agent. RL environments →
The set of processes (inference, environments, orchestration, trainer) that turns a policy into scored rollouts and scored rollouts into new weights. RL training systems →
Reinforcement learning from human feedback: train a reward model on human preference comparisons, then optimize a policy against it. Reward models, judges, and rubrics →
RLOO
REINFORCE with a leave-one-out baseline, where each rollout is compared with the mean reward of the other rollouts for the same prompt. Advantages and baselines →
RLVR
Reinforcement learning with verifiable rewards, where the reward comes from a program that checks the outcome, such as tests or an answer checker. Verifiable rewards (RLVR) →
Rollout
One complete attempt at a task, from the initial observation to termination. Rollouts and traces →
Rollout generation
The stage of RL training that runs the current policy on tasks through environments and collects completed rollouts. Rollout generation →
Router replay
Recording MoE expert selections at inference time and reusing them in the trainer so both compute the same routing. Trainer–inference mismatch →
Rubric
A prompt-specific list of weighted criteria, each graded by a judge, whose combined verdicts form the reward. Reward models, judges, and rubrics →
Sample
One training token sequence, built from one trainable branch of a trace. Rollouts and traces →
Sampling-mask replay
Recording which token IDs survived top-p or top-k at sampling time so the trainer renormalizes over the same set. Trainer–inference mismatch →
Sandbox
An isolated execution environment for an agent's actions, typically scoped to one rollout, whose runtime policy limits access to the host and other workloads. Sandboxes →
SDFT
Self-Distillation Fine-Tuning, which uses a demonstration-conditioned copy of the model as its own on-policy teacher. Teachers and privileged context →
Self-distillation
On-policy distillation in which the teacher is the student itself, conditioned on privileged context. Teachers and privileged context →
Self-play
Training a policy on interactions with copies of itself, such as proposer-solver pairs or games, so the curriculum moves with the policy. Self-play and multi-agent RL →
Sequence packing
Concatenating variable-length samples into fixed-length rows for the trainer while preventing attention across sample boundaries. Groups, batches, and packing →
Sequence-level mask
A trust region that drops a whole rollout when its average log-ratio between behavior and current policy exceeds a threshold; DeepSeek-V3.2 applies it only to negative-advantage rollouts. Importance sampling →
Sequence-level ratio
An importance ratio for a whole response, the product of its token ratios; GSPO uses its length-normalized geometric mean. Importance sampling →
Simulated user
A language model that plays the user in a conversational environment, which makes its replies part of the transition function. RL environments →
Speculative decoding
Generating with a cheap drafter whose proposals the policy verifies, so outputs keep the policy's sampling distribution. Inference-heavy workloads →
Training a model with cross-entropy to reproduce fixed demonstration tokens. Post-training methods →
Surrogate objective
A loss whose gradient approximates the policy gradient, such as a token-level importance-weighted loss that reweights next-token choices at recorded contexts. Importance sampling →
Target policy
The policy π_θ that the loss optimizes, which can differ from the behavior policy that sampled the data. Policies →
Task validation
Checking each candidate task before training so that its gold solution passes, a no-op fails, and repeated runs agree. Building environments at scale →
Taskset
The part of an environment that loads tasks and owns their data, task-specific tools, and scoring. RL environments →
Teacher
A model that either wrote the tokens a student trains on (SFT, hard distillation) or scores the student's own tokens (on-policy distillation). Post-training methods →
Tool call
A structured action encoded in sampled tokens that the harness parses and executes outside the model. Turns and tool calls →
Tool-use collapse
A feedback loop in which an agent stops calling helpful tools because early calls were associated with failure. Entropy collapse and exploration →
Trace
The stored record of a rollout: requests, sampled tokens, actions, observations, rewards, how it ended, and token authorship. Rollouts and traces →
Trainer–inference mismatch
The gap between the token probabilities the sampler used and those the trainer computes for the same tokens and weights. Trainer–inference mismatch →
Truncated importance sampling (TIS)
An off-policy correction that caps each importance weight at a constant and treats it as fixed. Importance sampling →
Truncated sampling
Sampling with top-p or top-k, which removes low-probability tokens from the behavior policy's support. Policies →
Truncation
A rollout ending forced by a turn, token, or time budget rather than by the task itself. Rollouts and traces →
Turn
One model generation; the observation that follows is not part of it. Turns and tool calls →
Value function (critic)
A learned model that predicts the expected return from a state, used as a baseline and for per-step advantages. Advantages and baselines →
Verifier
A program that decides whether a rollout's final answer or state satisfies the task. Verifiable rewards (RLVR) →