Harnesses

A harness is the program around a language model that runs the agent's side of an interaction. It builds each model request, exposes tools, parses sampled output into actions, executes or forwards those actions, returns observations, handles errors, and ends the rollout. Claude Code, Codex and a fifty-line tool loop are all harnesses.

The harness sits on the agent's side of the agent/environment boundary, so RL trains the model to act well inside one specific harness. Each harness choice changes which behavior training sees and reinforces.

Constructing the next model call

Changing context construction changes the policy input even when the model weights stay fixed.

Parts of a harness

Context construction. Each request contains a context assembled by the harness: the system prompt, the task, the interaction so far, and possibly a compacted summary, retrieved notes or a subagent's result. The model can only use what the harness puts there, and when the history outgrows the context window, the harness's compaction policy sets what survives.

The action language. Tools define what the agent can do and how it must express it. A single shell tool that takes a command string admits different behavior from separate tools for reading, editing and running files. Tool names, descriptions, argument types and error messages all change which actions the model is likely to produce. A tool is usually declared to the model as a JSON Schema:

{
  "name": "run_command",
  "description": "Run a shell command in the task sandbox and return its output.",
  "parameters": {
    "type": "object",
    "properties": {"command": {"type": "string"}},
    "required": ["command"]
  }
}

The name the model sees also depends on the harness. The same tool can appear as run_command, as shell_run_command in a harness that prefixes tools with their server name, or under an MCP namespace in another, and a model trained on one spelling has to map it to the others at deployment.

Parsing. The harness sets which sampled token sequences count as valid actions, how strictly to parse them, and whether to repair near-misses (turns and tool calls).

Errors. A malformed call can be rejected by the parser, a valid call can fail in the environment, and a successful call can return output the model misreads. The harness determines which of these become observations the model can react to and which end the rollout.

Stopping. The harness enforces budgets on turns, tokens and wall-clock time, and recognizes when the model has submitted a final answer. How the rollout ended affects scoring (rollouts and traces).

Worked example

Consider one coding task under two harnesses. Harness A gives the model a shell and up to fifty turns. Harness B asks for a unified diff in a single response. The repository and the tests are identical. A model that solves 60% of tasks under A might solve 25% under B, because B removes the ability to read files, run tests and revise.

The 35-point gap comes entirely from the harness. With the model held fixed, tool schemas, parsers, prompts, stopping rules and context policies all change the distribution of actions.

Training in the deployment harness

Two rules follow from the harness being part of the agent.

  • Comparisons between checkpoints hold the harness fixed. Comparing checkpoints under different prompts, tool formats or context policies confounds model improvement with interface improvement. "Which model plus harness is best" is a separate question from "did training improve the model", and it needs a separate measurement.
  • Training uses the harness that will be deployed. RL reinforces behavior that works inside the training harness, including its tool format, error style and compaction rule. When deployment differs, part of what was learned does not transfer, and the model may keep producing the training harness's formats.

Harness problems are also cheaper to fix than weight problems. A model may understand the task and still fail because a schema is awkward, the parser rejects a valid call, or the context omits a key file. Reading traces before training finds many of these.

Optimizing the prompt and harness

The harness is itself a set of parameters that can be optimized: the system prompt, tool descriptions and context policy. An outer loop can run rollouts, inspect the traces, propose a change, and keep changes that raise reward, with no gradient step on the weights. GEPA (Agrawal et al.) does this for prompts by having a model reflect on traces in natural language. Across six tasks it outperformed GRPO by 6% on average and by up to 20%, while using up to 35× fewer rollouts.

The two kinds of optimization combine. Ziems et al. found that multi-module GRPO composed with automatic prompt optimization raised accuracy by 11% on average over the unoptimized post-trained model, and by 5% over prompt optimization alone. A prompt edit can fix an interface problem that RL would need many gradient steps to work around. At low to medium compute, optimizing the prompt or harness first and then running RL inside the optimized harness can beat spending the same budget on RL alone. The order matters for the previous section's rule: RL trains the model for the harness it runs in, so the optimized prompt becomes part of the deployment harness and stays fixed afterwards.

Effects on training data

The harness sets which tokens exist in each request, and so sets what training data looks like. Tool results appear in context between turns and condition later tokens without a policy-gradient term of their own (loss masks). When the harness rewrites earlier context, the next request no longer extends the previous token sequence, and the trainer starts a new sample at that point (token-level rendering). Time spent waiting on tools occupies a rollout slot without generating tokens, which shapes rollout generation.