Building environments at scale
Environment scaling is the work of producing agent tasks by the thousands, each with a reproducible starting state and a scorer, and checking every one before training uses it. It sets the ceiling on agent RL: a policy learns only the skills its tasks exercise, and a broken task either wastes its rollouts or, when it rewards a wrong answer, trains the policy toward that answer.
Sources of tasks
Agent tasks come from five kinds of source, which trade realism against volume:
- Mined repository history. A merged pull request that fixes an issue gives a starting commit, a problem statement, and tests that fail before the fix and pass after it. R2E-Gym extends this to commits with no linked issue by generating tests and back-translating the change into a problem statement, collecting more than 8.1K tasks.
- Synthesized bugs. SWE-smith breaks working code on purpose: a model rewrites functions, AST transforms mutate them, validated bugs are combined, and merged fixes are reverted. A candidate survives only if it breaks tests that passed before. The pipeline produced 50K tasks from 128 repositories, with yields per method from 33.8% to 96.9%.
- Procedural generation. Endless Terminals has a model write a terminal task, build and validate its container, write completion tests, and filter for solvability, producing 3,255 tasks. Plain PPO on them lifts Qwen2.5-7B from 10.7% to 53.3% on its dev set.
- Synthetic tool environments. AgentScaler groups more than 30,000 APIs into over 1,000 domains and treats each tool call as a read or write on a domain database, so a rollout is checked against the database state it produces. Kimi K2 pairs over 3,000 existing MCP tools with more than 20,000 synthetic ones.
- Self-play. The policy writes its own tasks. Self-play SWE-RL has one model inject bugs into open-source repositories and then fix them, with no human-written issues or tests (self-play).
Every source produces candidates, and a candidate is not yet a task. Mined tasks inherit broken builds and flaky tests from their repositories, and synthesized tasks inherit the generator's mistakes.
Worked example
Prime Intellect's Scaling Agentic RL post reports what validation removed from five public SWE datasets. Each gold patch ran through the full scoring path in fresh sandboxes, failures were retried up to 10 times, a second independent pass caught rows that flipped, and no-edit passes dropped tasks that score 1.0 without any change.
| Dataset | Raw tasks | Verified | Kept |
|---|---|---|---|
| R2E-Gym (subset) | 4,578 | 4,522 | 98.8% |
| SWE-Lego (real data) | 4,432 | 4,323 | 97.5% |
| Scale-SWE | 20,181 | 17,202 | 85.2% |
| Multi-SWE | 4,703 | 2,232 | 47.5% |
| SWE-rebench-V2 | 32,079 | 6,275 | 19.6% |
The losses have different causes. R2E-Gym's 56 drops are mostly network- and timing-sensitive tests. Multi-SWE lost more than half its rows to two gold passes and a no-edit filter. SWE-rebench-V2 dropped whole languages whose images were broken, removed flaky rows, and had its problem statements scrubbed of inline links to the GitHub issue or pull request. A run on the raw SWE-rebench-V2 release would have spent four of every five rollouts on tasks that validation rejects.
The validation pipeline
Each check targets one way a task misreports success:
- The gold solution passes. A task whose reference fix fails its own scorer has a broken image, a missing dependency, or a wrong reference, and every rollout on it scores zero.
- The no-op fails. A task that scores 1.0 with zero edits rewards doing nothing, and any rollout that stops early is reinforced.
- Reruns agree. A test that passes on some runs and fails on others turns the reward into a coin flip. Retrying separates flaky from deterministically broken.
- Network dependencies are removed. A test that downloads a package or calls a service passes during validation with network access and fails inside an offline training sandbox.
Together these checks pin the scorer at two known points, the gold solution and the empty one.
False negatives from implementation-specific tests
A gold solution passing proves the tests accept one fix, the one they were written against. Tests often encode incidental details of that fix: a helper function's name, the exact text of an error message, the order of dictionary keys in a printed result. A different correct fix fails them. The gold check cannot catch this, because the gold patch is the implementation the tests describe.
False negatives teach the policy to imitate the reference implementation's surface instead of solving the problem. Rewriting tests to check behavior removes the cause, at a cost per task. Hybrid verification adds a second opinion by pairing test execution with a model verifier. An agentic verifier goes further, reading the patch and running its own probes before scoring (reward models and judges).
Keeping grading material out of reach
In most SWE tasks the agent shares a machine with the grading machinery, and a policy under RL pressure finds any leak. The runtime protections belong to the sandbox: hidden tests arrive only after the agent stops, or grading runs elsewhere. The task data needs its own scrubbing, because the upstream fix is recorded in the repository and the issue tracker. .git history is truncated at the starting commit, and problem statements are stripped of links to the fixing issue or pull request (verifiable rewards).
Difficulty and mixing
A verified task can still be useless for training. A task the current policy always solves or always fails gives a group of rollouts zero advantage, so pools are filtered by the policy's pass rate before and during training (cold start). In a recipe study on the TravelPlanner benchmark, Wu et al. found that about 1K training tasks with a balanced mix of difficulties beat larger sets on out-of-distribution tests.
A run usually mixes environments: SWE tasks, terminal tasks, search tasks. Their rollouts differ by orders of magnitude in length and wall-clock time, so the mixture ratio sets both what is learned and how batches fill (batching and packing).
Environment stability
Infrastructure failures change the training signal. Wu et al. injected random tool-execution failures while training a 3B model and evaluated it without them. Test performance held up at failure rates up to 5%, and at 10% training stability and convergence speed dropped clearly. A sandbox that times out under load, or a tool server that returns errors, shows up as reward noise in every environment that depends on it. Per-environment error and timeout rates therefore belong on the training dashboard next to reward, and a failed rollout is excluded instead of scored zero (rollouts and traces).
Hubs and reuse
A task packaged with its image, tools, and scorer runs the same way for evaluation and for training. The benchmark number and the training reward then come from one scoring path, and a fix to a flaky test reaches both. Environment hubs extend this across teams: a published, versioned environment lets one lab train on another's tasks, and lets an evaluation name the exact version it ran (evaluating agents).