Evaluating agents
Evaluating an agent means measuring its behavior on a fixed set of tasks under a fixed procedure, so that results can be compared across checkpoints, models or harnesses. Training reward says whether optimization is succeeding against the reward. Evaluation asks what the policy can now do, and for agents the conditions it holds fixed include a harness, a runtime and budgets, each of which can move a score as much as a better checkpoint.
The evaluation protocol
Two scores are comparable only if they were produced the same way. The protocol of an agent evaluation includes:
- the tasks and the split they come from
- the harness, its version and its tool set
- the runtime image and its network policy
- sampling settings, which are part of the behavior policy
- budgets: token limits, turn limits and timeouts
- the number of rollouts per task
- the scorer, including the judge model if one is used
- for conversational tasks, the user simulator's model and prompt
A higher token limit lets more tasks finish, and a different harness changes what the model sees, so a report states whether it used a standard harness, which compares checkpoints, or the product harness, which measures what users get. A simulated user is a second sampled policy, so its randomness adds variance of its own, and changing its model can move the score with no change to the agent. Success belongs next to tokens, turns, wall time and cost, because a policy can raise its success rate by spending more of each.
pass@k and its unbiased estimator
pass@k is the probability that at least one of independent attempts at a task succeeds, averaged over tasks. pass@1 measures reliability, and larger measures coverage, meaning whether the policy can solve the task at all.
The standard estimator, from the HumanEval paper, draws samples per task, counts the that succeed, and computes the probability that a random subset of contains at least one success:
The fraction is the probability that all chosen samples are failures. The estimator is unbiased, while the plug-in estimate is biased low. With and :
- pass@1
- pass@5
- the plug-in estimate for is
Binomial coefficients overflow for large , so implementations use a product form:
import numpy as np
def pass_at_k(n: int, c: int, k: int) -> float:
if n - c < k:
return 1.0
return 1.0 - np.prod(1.0 - k / np.arange(n - c + 1, n + 1))The complementary pass^k, introduced with τ-bench, is the probability that all attempts succeed, estimated as . With of , pass^2 . It measures the consistency that matters for an agent users run many times. pass@k also differs from best-of-k: pass@k assumes an oracle that recognizes the success, while best-of-k needs a selector, such as a reward model or the agent's own tests, to pick it.
The two ends of the axis can move differently under RL. Yue et al. found RLVR models beating their base models at small while the base models reached higher pass@k at large , a result that longer training and stricter success criteria complicate (RL scaling). Reporting several values of shows which end moved.
Held-out tasks and environments
A held-out split of the training distribution tests whether gains transfer to instances that never provided gradient, and a held-out environment, with different tools, state or verifier, tests a stronger form of transfer. Training reward that rises while held-out performance stays flat points to memorized artifacts or reward hacking.
A split is only as independent as its construction. Tasks from the same repository, templates filled with different values, or near-duplicate problems leak across a random split, so splitting by repository, source or template measures transfer more faithfully. The suite, its primary metrics and its sampling budget should be fixed before the run, so the headline comparison is not selected afterward.
Contamination
A benchmark whose tasks appeared in pretraining data measures recall. Wu et al. gave Qwen2.5-Math-7B the first 60% of each MATH-500 problem and found it completed 54.6% of them word for word, against 0% on the newer LiveMathBench. Prefix-completion tests, benchmarks released after the model's data cutoff, and private task sets are the available checks.
What agent benchmarks fix
Public agent benchmarks differ in which parts of the protocol they fix:
| Benchmark | Measures | Fixes | Leaves open |
|---|---|---|---|
| SWE-bench Verified | 500 Python repository issues | human-screened specs and tests | public repositories, so contamination |
| SWE-Bench Pro | 1,865 multi-file tasks from 41 repositories | held-out and commercial repositories | harness and budget choice |
| Terminal-Bench 2.0 | 89 command-line tasks, each with its own environment | human-written solutions and tests | at 89 tasks, a standard error near 5 points at 50% success (binomial estimate) |
| τ²-bench | support tasks where agent and user both act on shared state | task generator and constrained user simulator | user-simulator variance |
| BrowseComp | 1,266 hard-to-find facts on the web | short answers checked against a reference | the live web changes and answers can leak |
| METR time horizon | human task length at 50% agent success | a unit comparable across models | depends on the task suite and human baselines |
Variance and confidence intervals
An eval score is an estimate, and its uncertainty is often larger than the difference being reported. Treating the tasks as a sample from a larger population, as Miller recommends, gives a standard error for the mean over tasks. For tasks with one binary attempt each and a success rate of :
A 95% confidence interval is about , or .
More rollouts per task help less than expected. With rollouts per task, the variance of the mean splits into a part from differences between tasks and a part from sampling within each task:
where is the policy's true success rate on task . Suppose and , which add up to the total above. Going from to shrinks the standard error from to , and only reaches . Most of the uncertainty comes from which tasks are in the set, and only more tasks reduce it. Rollouts per task are still needed for pass@k and per-task difficulty.
Two checkpoints evaluated on the same tasks call for a paired comparison: compute the per-task difference and its standard error . Task difficulty affects both checkpoints alike and cancels. If a new checkpoint goes from 80 to 90 successes on 200 tasks by fixing ten tasks and breaking none, the improvement is 0.05 with a paired standard error of about 0.015, a 95% interval of roughly , even though each score's own interval is wider than 0.05.
Reading beneath the mean
An aggregate hides which tasks changed. Per-environment results belong in the report next to any macro average, and difficulty slices show whether a policy improved only where the base model already succeeded occasionally. The fractions of tasks where every rollout succeeds or every rollout fails show whether the frontier of partially solved tasks is moving, which connects evaluation to difficulty filtering. Error and truncation rates belong in the report, because a crash or a slow sandbox is not a wrong answer.
A small, stable set of traces reviewed across checkpoints separates failures that look identical in the score: a reward shortcut, a harness that hid needed information, an environment timeout, a verifier that rejected valid work, or an improvement bought with far more compute.