Rollouts and traces
A rollout is one attempt at a task, from the initial observation to termination, with the agent and environment alternating actions and observations. A trace is the stored record of that rollout: every model request and the tokens sampled in response, the parsed actions, the observations returned, the rewards, and how the rollout ended.
Rollouts are the unit of experience in agent RL, and the trace is the only form in which that experience reaches the trainer. Whatever the trace fails to record, such as the exact token IDs or why the rollout stopped, cannot be recovered when the training batch is built.
A sampled rollout
Rollouts, episodes, branches, and samples
The environment supplies the task and initial observation, and the model generates a turn. If the turn ends in a tool call, the environment executes it and the result becomes the next observation. The loop repeats until something ends it. Four terms describe what that loop produces.
- A rollout is one agent's attempt at one task.
- An episode is the container for rollouts that belong together: a solver and its judge, the players in a game, or several agents' rollouts in one multi-agent run.
- A branch is one root-to-leaf path through a trace's message graph. A linear tool loop has one branch; compaction and subagents add more.
- A sample is one training token sequence. Each trainable branch becomes one sample.
How rollouts end
| Ending | Example | What it says about the policy |
|---|---|---|
| Submitted | the agent gives a final answer | the answer can be scored normally |
| Task-detected | the environment sees the goal reached, or an unrecoverable state | an outcome of the task |
| Truncated | turn, token or time budget exhausted | the attempt was cut off, not finished |
| Length-cut | one generation hits its token limit mid-output | often a runaway or malformed turn |
| Error | a sandbox crash, a tool server failure, a lost connection | usually nothing about the policy |
Only the first two are outcomes of the task. An infrastructure error says nothing about the policy, and a truncation says the policy did not finish within the budget, which differs from failing.
Worked example
A task is sampled 8 times. Three rollouts submit a passing fix, two submit a failing one, one is truncated at the turn limit with the bug unfixed, and two die when their sandbox host goes down.
Scoring every ending as reward 0 gives rewards and a group mean of . The two errored rollouts receive advantage , so the update lowers the probability of whatever tokens the model had produced before the host failed. The passing rollouts receive , inflated by two failures that were never attempts.
Dropping the errored rollouts leaves with mean . The passing rollouts receive , the failing and truncated ones , and the errored ones contribute no gradient. The truncated rollout is scored by the task's normal scorer on its final state; it scored 0 here because the bug was still present, and a truncated rollout that had already fixed the bug would score 1.
The example suggests a default: drop errored rollouts from the batch before computing advantages, score truncations with the normal scorer, and log the ending type of every rollout so that error and truncation rates can be tracked. Penalizing truncation beyond its score is a separate decision, appropriate when efficiency is part of the task.
What a trace records
A trace serves several readers.
- The scorer needs the actions and final state.
- The trainer needs token IDs, a flag marking which tokens the policy sampled, the behavior policy's per-token log-probabilities, rewards, and later advantages.
- A researcher needs enough structure to see why the agent succeeded or failed.
- An operator needs timing, tool errors and resource use.
All of these views should come from one record. Reconstructing training data from chat logs after the fact loses the exact token IDs, and rendering the same messages again can produce different tokens (token-level rendering). The sampled-token flags become the loss mask, and the behavior log-probabilities are what importance sampling compares against once the trainer's policy has moved on.
The message graph
A linear transcript is enough for a simple tool loop. Agents that compact their context or call subagents break that shape. After compaction, later requests see a summary instead of the earlier turns. A subagent runs its own sequence of requests and returns only a result to its parent. A flat list of messages cannot represent both the subagent's history and the parent's later view without losing which request saw what.
A trace stores each event once
Storing the trace as a graph of messages keeps shared history once and makes each root-to-leaf path a branch, a contiguous token sequence that can be trained as its own sample. The transcript a person reads is one branch through that graph. When a metric moves during training, the branches show which behavior changed: shorter solutions, more test runs, a new way of editing files, or a hack.