RL training systems

An RL training system is the set of processes that turns a policy into scored rollouts and scored rollouts into new weights: an inference engine samples tokens, environments run and score each rollout, an orchestrator groups finished rollouts into batches, and a trainer computes the update and publishes the new weights. In agent RL, every hand-off between these processes can change the data an update is computed on, so the system design is part of the algorithm.

Worked example

Take a small run: a 0.6B-parameter model, already warmed up with SFT, learning to reverse short sentences character by character. The reward gives partial credit, the similarity between the model's answer and the true reversal. One GPU serves inference, one GPU trains, and the orchestrator runs on the CPU. Each step trains on 128 rollouts, 8 tasks with 16 rollouts each, sampled at temperature 1.0 with completions capped at 128 tokens.

  1. Dispatch. The orchestrator picks 8 tasks and sends 16 run requests per task. Each request carries a group ID shared by the 16 rollouts of its task and the policy version current at dispatch. In the default asynchronous mode this happens while the trainer is still working on the previous batch, so the rollouts come from weights one version behind the ones they will update (async RL).
  2. Generation. The inference engine samples each completion and returns its token IDs along with the log-probability of every sampled token. For the task "In 1891 the community inaugurated its own cemetery", the prompt renders to about 50 tokens.
  3. Scoring. The environment extracts the text inside the answer tags and computes its matching-subsequence similarity ratio (Python's difflib.SequenceMatcher) to yretemec nwo sti detaruguani ytinummoc eht 1981 nI. A perfect reversal scores 1.0, one dropped letter 0.99, an unreversed year 0.98, and the right words in reverse order but spelled forwards 0.40.
  4. Advantage. When all 16 rollouts of a group have arrived, the orchestrator subtracts the group mean. If the group averages 0.62, the 0.99 answer gets advantage +0.37 on each of its sampled tokens and the 0.40 answer gets −0.22. A group whose 16 rewards are identical has zero advantage everywhere, so it is dropped and replaced, and the step still trains on 128 rollouts with signal (GRPO).
  5. Packing. Each rollout becomes one training sample of about 150 tokens: the token IDs, a mask marking the sampled tokens, their recorded log-probabilities, the sampling temperature, and the advantages. The 128 samples, about 19,000 tokens, are packed 13 to a row into rows of 2,048 tokens, which gives 10 rows (batching and packing).
  6. Trainer step. For each row, the trainer runs a forward pass, computes the log-probability of every sampled token at the sampling temperature, and forms the ratio ρt\rho_t against the recorded value. It accumulates gradients of the loss over all 10 rows, then takes one optimizer step.
  7. Weight broadcast. The trainer sends the new weights, about 1.2 GB for 0.6B parameters in bf16, to the inference engine, which loads them and advances its policy version (weight broadcast). Rollouts already in flight continue under the new weights.

The run takes 20 such steps. Average evaluation reward is about 0.05 for the base model and about 0.8 after the SFT warm-up and the 20 RL steps.

Four jobs, four resource profiles

The steps above fall into four jobs with different resource needs. Generation decodes variable-length sequences and is usually bound by memory bandwidth, and agent rollouts add waits on tools and sandboxes between model calls. Scoring takes milliseconds for a verifier and a full inference request for a judge. Training runs dense, predictable tensor operations and benefits from tightly connected accelerators. Orchestration sits between the other three: it dispatches tasks, closes groups, computes advantages, forms batches, and tracks which weights inference holds.

In the example, inference and training each hold one GPU and everything else runs on the CPU. At scale the same split becomes separate node pools, and the jobs can overlap in time (scaling agent RL). When the jobs run in lockstep, each pool sits idle while it waits for the slowest one.

Who owns each fact

A sample reaching the trainer carries facts produced by different processes, and each fact has one source.

FactSource
Sampled token IDs and their log-probabilitiesInference engine
Which tokens the policy wroteRenderer, recorded in the trace
Reward and termination reasonEnvironment
Group membership and advantageOrchestrator
Policy versions that sampled the tokensOrchestrator
Parameter updateTrainer

Recomputing one of these facts somewhere else changes the algorithm. Re-tokenizing the transcript instead of keeping the sampled IDs makes the trainer score a context the model never saw (token-level rendering). Recomputing behavior probabilities in the trainer instead of recording them at sampling time hides the trainer–inference mismatch. Dropping the version record removes the only measure of policy lag. Jobs can move, with scoring inside the environment worker or orchestration inside the trainer process, as long as each fact keeps one source.

Timing settings

How many rollouts are in flight (rollout generation), when a group closes, how often weights reach inference, and how stale a rollout may be before it is dropped all change the distribution an update is computed on. In the example, a rollout still queued when the trainer has moved more than eight versions past the weights that sampled it is discarded instead of trained on (async RL). Two runs with the same loss and hyperparameters therefore train on different data when their systems overlap work differently.