Inference-heavy workloads
An inference-heavy workload is an RL run in which producing experience costs far more than learning from it: most GPU time goes to generating rollouts, and the trainer spends much of its time waiting for data. Agent RL is usually in this regime, because long generations, many turns, groups of attempts per task, large tool outputs, and model judges all add inference, while the update is one forward and backward pass over the sampled tokens.
Counting the work around a training token
A sampled token can be computed several times before it contributes to an update. The policy generates it. Later turns may prefill it again. The trainer evaluates its probability for the loss. A teacher or reference model may score it once more. A rough accounting:
The terms use different hardware and are not interchangeable, but the decomposition shows what a run pays for. An on-policy distillation run adds teacher inference. A run scored only by verifiers pays no judge inference, and its goes mostly to sandboxes.
Worked example
Raw generated tokens overstate throughput, because uniform-reward groups, invalid traces, stale samples, and infrastructure failures are all filtered out before the update. Usable experience multiplies through the losses:
dispatched rollouts
× completion rate (no infra failure)
× scoring rate (a reward was produced)
× nonzero-signal rate (group rewards vary)
× freshness rate (within the staleness bound)
= usable experienceWith 1,000 dispatched rollouts and rates of 0.95, 0.98, 0.60, and 0.90, about 500 reach the trainer. Raising the nonzero-signal rate from 0.60 to 0.80 with better task selection yields about 670, as much as a one-third increase in decode throughput would, without touching the serving stack. The freshness rate depends on the asynchronous RL staleness bound, and all four rates drift as the policy's behavior changes the workload. Hiding tool waits behind concurrent rollouts is an admission-control problem covered in rollout generation.
Prefill and decode
Transformer inference has two phases. Prefill processes the prompt in parallel and builds the KV cache, and is compute-bound. Decode generates tokens one at a time, rereading the weights and cache for each, and is usually bound by memory bandwidth.
Agent traffic shifts the balance toward prefill. A coding agent turn that receives a 3,000-token test log and replies with a 150-token command prefills 3,000 new tokens and decodes 150, a 20:1 ratio for that turn, even with prefix reuse. Without reuse it re-prefills the entire conversation, and the ratio grows every turn. Prime Intellect's RL at 1T scale report found model and environment pairs with prefill:decode token ratios as high as 4:1. A reasoning task that emits one long chain of thought sits at the other extreme, dominated by decode.
Several techniques remove repeated prefill, and they combine. Exact prefix continuity across turns avoids re-encoding history (token-level rendering). Prefix caching lets a group share the task prompt. Compaction shortens the visible context at the cost of a new training sample.
Prefill–decode disaggregation
When the two phases share a GPU, a burst of long prefills stalls every request decoding on it, and one large tool output can add latency to many other rollouts. Prefill–decode (PD) disaggregation runs the phases on separate pools: prefill workers build the KV cache and transfer it to decode workers, which continue generation. Each pool can then be sized and configured for its phase. DistServe and Splitwise introduced the design for serving.
Moving KV caches across the network adds bandwidth, latency, and scheduling work, and both pools must be kept balanced. PD disaggregation pays off for a stable, measured imbalance, such as large observations and short replies at a scale where turn latency limits rollout throughput. For small models or single-turn tasks the transfer costs more than it saves.
Speculative decoding in rollouts
Speculative decoding lets a cheap drafter propose several tokens that the policy then verifies in one forward pass, accepting or replacing each by a rejection rule (Leviathan et al., 2023). The rule makes it lossless with respect to the sampled distribution: tokens are distributed as if the policy had sampled them one at a time, provided verification applies the same temperature and truncation. Rollouts therefore stay on-policy, and importance ratios stay valid as long as the recorded log-probabilities are the policy's from the verification pass. Recording the drafter's probabilities would change the optimization target.
Two results show the size of the gain in RL.
- Iso et al. (2026) integrated speculative decoding into NeMo-RL and measured 1.8× rollout throughput for an 8B reasoning workload under synchronous RL. The gain is smaller under asynchronous RL, which already hides part of generation behind training, and their simulator projects up to 2.5× end-to-end speedup at 235B scale when the two are combined.
- DAS (Shao et al., 2025) targets the long tail of rollout lengths. It builds its drafter from recent rollouts with a suffix tree and gives long generations more speculative budget, cutting rollout time by up to 50% with identical training curves.
The speedup depends on how closely the drafter matches the current policy. DAS found that rollouts from earlier checkpoints make worse drafts than recent ones. Iso et al. found that a drafter trained on in-domain data for the RL prompts beat one trained on chat data (1.77× against 1.51×) and gained almost nothing from further updates during RL (1.78×), so online drafter updates mainly protect a poorly matched drafter.