Scaling agent RL

Scaling agent RL means growing a training run across many accelerators and environment workers while keeping every sample attributed to the task, policy version, and environment state that produced it. An agent run joins three workloads with different resource needs: inference generates model turns, environments execute tools and hold state, and training computes gradient updates. Each is sized on its own, and the balance among them shifts as the policy learns.

Inference in agent loops

Agent rollouts alternate generation with environment waits: a short command, a build, another turn, a verifier. Latency per turn sets how long each rollout takes, and a large pool of concurrent rollouts keeps the engine busy while individual agents wait (rollout generation). The prefill and decode costs of that traffic are covered in inference-heavy workloads.

Environment capacity

A math verifier needs little more than a CPU process. A coding environment needs sandboxes, container images, and storage. A browser task depends on network services. A model judge is another inference workload with its own batching. Image storage grows with task diversity: Prime Intellect's agentic RL environments post describes about 135,000 prebuilt task images for software-engineering, terminal, and search tasks, kept in a registry co-located with the sandboxes that start them.

The environment pool is sized from observed service time. Too few runtimes leave the model waiting for observations. Too many create startup storms, exhaust storage, or overload shared services.

Worked example

Suppose an inference engine runs most efficiently with 64 sequences decoding at once. An average agent turn spends 2 seconds generating and 8 seconds waiting on a tool, so each rollout is decoding 20% of the time. Keeping 64 sequences decoding needs 64/0.2=32064 / 0.2 = 320 rollouts in flight per engine, and 320 live sandboxes to hold their state. Eight engines need 2,560.

The ratio moves with the policy. If later checkpoints write longer reasoning before each command, say 4 seconds of generation per 8 seconds of tool time, each rollout decodes a third of the time and 192 rollouts per engine suffice. The 320 rollouts sized at step 0 then want about 107 decoding slots against 64, so turns queue and every rollout slows down.

Training layout

Gradient computation exchanges parameters, activations, and optimizer state every step, and wants high-bandwidth links. Fully sharded data parallelism (FSDP) partitions parameters and optimizer state, tensor and expert parallelism split computation inside a layer, and context parallelism divides long sequences across devices. Agent traces reach very long contexts, so sequence length can shape the trainer layout as much as parameter count.

Colocated or disaggregated

The main placement question is whether training and inference share GPUs.

Colocated systems put both on the same devices and alternate: generate, switch engines and train, switch back. Weights stay local, so there is no network broadcast. Neither phase runs while the other does, and both engines must fit in the same memory. HybridFlow (verl) supports this arrangement and reshards the actor between its training and generation layouts.

Disaggregated systems give training and inference separate pools. Each scales independently with the hardware and parallelism that suits it, and they run at the same time, which is what makes asynchronous overlap possible. In exchange, rollouts move toward the trainer, weights move back through weight broadcast, and every sample needs its policy version tracked.

Long, uneven rollouts favor disaggregation. The GLM-4.5 report used both modes in one framework: colocated, synchronous training for math and code reasoning, and disaggregated, asynchronous training for agent tasks such as software engineering, whose long environment interactions would otherwise stall synchronous steps.

Workload drift during training

The policy changes the workload. Early in training a policy may produce short failed attempts. As it improves it may use more tools, reach deeper task states, and produce longer successful traces, so sequence lengths, turn counts, and tool latencies drift. A capacity ratio chosen at step 0 can leave the wrong resource saturated by step 500: decode capacity, sandboxes, or trainer throughput.

Drift also interacts with the staleness bound. The long successful traces a run is trying to reinforce are the ones most likely to exceed it, so an undersized inference pool selects against the behavior being learned. Completion and freshness rates tracked by rollout length, reward, and environment catch this.

Telemetry that follows a rollout across boundaries shows where time goes:

queue wait → prefill → decode → tool wait → verify → batch wait → train

A long rollout can reflect long generation, a slow build, a stuck worker, or a group waiting on its slowest member, and each cause has a different fix.

Decentralized rollout generation

Rollout generation has a narrow interface. A worker receives a policy and a task, and returns a trace with tokens, probabilities, and rewards. That makes sampling much easier to spread across machines or sites than gradient computation, and it moves the trust boundary: the trainer has to trust that a worker used the expected model, renderer, sampling settings, and environment. On a managed cluster, such as the 64-node cluster that trained INTELLECT-3 (Prime Intellect, 2025), that trust comes from operational control.

Open networks need verification instead. INTELLECT-2 (Prime Intellect, 2025) trained a 32B model with asynchronous RL on permissionless inference workers and checked every submitted rollout with TOPLOC plus sampling and sanity checks. TOPLOC hashes intermediate activations into a proof of 258 bytes per 32 generated tokens, and detected changes to the model, prompt, or precision in every test. SAPO goes further: each node trains its own policy and shares rollouts with the others instead of gradients. It reported cumulative reward gains of up to 94% in controlled experiments and ran on thousands of community nodes.