Token-level rendering

Token-level rendering is the practice of building each model prompt in a multi-turn agent rollout directly from token IDs, so that the tokens the model sampled on earlier turns appear in later prompts unchanged. A renderer is the component that converts between the message view that harnesses and environments use and the token view that the model samples from and the trainer learns on. For single-turn tasks a renderer is just a chat template. For agents it determines whether the training data matches the contexts the model generated in.

Why the prompt must extend the sampled tokens

The behavior log-probability recorded for a sampled token was computed under one specific token prefix:

μ(yt∣x,y<t).\mu(y_t \mid x, y_{<t}).

The trainer later evaluates the same token under its own weights and forms a ratio against μ\mu (importance sampling). If the trainer's copy of the context differs from the sampler's copy by even one token, it scores a different conditional probability, and the ratio compares two unrelated numbers.

The naive agent loop creates this problem on every turn. It parses the model's reply into a structured message, runs the tool, appends the result as another message, and renders the whole conversation again with the chat template. The rendered text looks the same, but the token IDs often differ.

How token identity is lost

Several ordinary steps can change tokens while preserving apparent meaning:

  • Parsing and re-serializing tool calls. The model samples {"verbose": false}. The harness parses it into a Python dict, and the template re-serializes it as {'verbose': False} or with different whitespace and key order.
  • Different merges. A BPE tokenizer can encode the same string differently depending on neighboring bytes, so re-tokenizing a region after its surroundings changed can produce different IDs.
  • Template history rules. Many chat templates drop or rewrite earlier reasoning blocks when rendering history, so a full re-render removes tokens the model conditioned on.
  • Structural tokens. Assistant openers, tool-call delimiters, reasoning markers and end-of-turn tokens depend on the model family and on position. A truncated turn may have no end-of-turn token at all, while the template adds one on re-render.

Each of these yields a prompt the model never saw, and a trainer that trusts it computes a log-probability for a context that did not exist.

Render, parse, bridge

A token-level renderer separates three operations that the naive loop collapses into one:

  1. Render messages into token IDs for the first prompt, with each token attributed to its source message.
  2. Parse the sampled completion IDs into a structured reply (content, reasoning, tool calls) for the harness to act on, without touching the IDs themselves.
  3. Bridge to the next turn by appending only the new environment messages and the next assistant opener to the previous prompt plus completion.

The bridge has a precise contract. If pp is the previous prompt and cc the sampled completion, the bridged sequence BB must satisfy B[:∣p∣+∣c∣]=p∥cB_{[:|p|+|c|]} = p \mathbin\Vert c. When the renderer cannot guarantee that, for example because the model stopped inside an unfinished reasoning block or the new messages include assistant content that would be re-tokenized, it refuses and the caller falls back to a full render.

The per-token attribution from rendering is the provenance that loss masks use to select which tokens are trained.

Prefix breaks

Some breaks are accidental and some are deliberate. Compaction replaces old context with a summary. A template may drop old reasoning at a new user message. A subagent begins from a fresh prompt. In each case the new prompt does not extend the old token stream.

The correct representation is to start a new training sample at the break. The earlier tokens remain one contiguous sample with correct probabilities, and the new prompt becomes fixed context for the next sample. Hiding the break would assert a factorization of the sequence that generation never used.

Breaks cost compute as well as correctness. If every turn is re-rendered, each turn's prompt becomes its own sample, and a five-turn rollout trains on overlapping prefixes of one to five turns.

Exact token prefixes across turns

With five equal-length turns, re-encoding every prefix processes 1 + 2 + 3 + 4 + 5 turn-lengths. Preserving the prefix processes 5.

A high rate of breaks in a run can mean long tasks that need compaction, or a harness that re-renders when it does not need to. Counting samples per rollout separates the two. The same exact-prefix invariant is what lets an inference engine reuse its cached state across turns (KV cache).

Reasoning across turns

Reasoning models write a thinking block before each reply, and chat templates differ in whether that thinking stays in the context of later turns. Qwen3's template keeps reasoning only for assistant turns after the most recent user message. Within one tool-calling cycle the reasoning stays and each turn bridges, but when a new user message arrives, earlier reasoning disappears from the rendered history and the prefix breaks.

Keeping reasoning across user messages, sometimes called interleaved thinking, lets a long conversation stay one sample at the cost of a longer context. Dropping it starts a new sample at every user message. The choice also changes what the policy conditions on at each turn, so training uses the retention rule that deployment uses (harnesses).