Self-play and multi-agent RL

Self-play trains a policy on interactions with copies of itself, so the training distribution moves as the policy improves. Multi-agent RL is the broader setting in which one episode contains several agents, each producing its own trace and its own reward. For LLM agents, the common forms are a proposer that writes tasks for a solver, two players in a game, and a solver paired with a critic or judge.

The appeal is an automatic curriculum. A fixed taskset eventually becomes too easy, and tasks that are too hard give no signal because every rollout fails (cold start). A proposer that learns to write tasks near the solver's current ability keeps outcomes mixed within each group, which group-relative methods such as GRPO need.

Proposer and solver

In a proposer-solver setup, one agent generates a problem and several solver rollouts attempt it. The solvers are rewarded for solving. The proposer is rewarded for what its problem does to the solvers.

A good proposer reward peaks on problems that some solvers solve and others fail. One choice is the learnability score 4p(1−p)4p(1-p), where pp is the fraction of solvers that succeed. With four solvers per problem:

Solvers who succeed (of 4)pp4p(1−p)4p(1-p)
00.000.00
10.250.75
20.501.00
30.750.75
41.000.00

A problem every solver cracks and a problem none can crack both earn zero, and both would have given the solver group zero advantage anyway. The proposer is paid for creating the variance that solver training needs.

Published systems vary the shape of this reward and the source of the answer:

  • R-Zero rewards its Challenger with 1−2∣p^−12∣1 - 2\lvert\hat p - \tfrac12\rvert, a tent that also peaks at 50%, and takes the answer from the Solver's own majority vote, so no external label exists. It raises Qwen3-4B-Base by 6.49 points on math benchmarks starting from no data.
  • SPICE grounds the proposer in a document corpus. The Challenger writes a question and gold answer from a document the Reasoner never sees, which keeps answers checkable without a hand-built verifier. Averaged over four base models, it gains 8.9 points on math and 9.8 on general reasoning.
  • Self-play SWE-RL (SSR) applies the idea to coding agents in open-source repositories. One policy injects a bug, writing a bug patch, a test script, and a patch that weakens the tests that would reveal it. The same policy then repairs the bug, seeing the reversed test-weakening patch as its specification instead of an issue. A proposal that fails validation (for example, no test fails after injection) earns −1-1. SSR improves by 10.4 points on SWE-bench Verified and 7.8 on SWE-Bench Pro, and stays ahead of a baseline trained on human-written issues and tests throughout training.

The answer source bounds the reward's accuracy. A majority-vote label is wrong whenever most solvers agree on a wrong answer, while a gold answer taken from a document or a test that executes does not depend on the solvers at all.

Comparison groups for advantages

Group-relative baselines assume the rollouts being compared are exchangeable rollouts of the same task. Multi-agent episodes break that assumption in three ways.

Different problems. Solver rollouts on problem A are not comparable with rollouts on problem B, which may be much harder. Solvers are compared only with other rollouts on the same proposed problem. If three solvers on one problem get rewards 1,1,01, 1, 0, the baseline is 2/32/3 and the advantages are (+1/3,+1/3,−2/3)(+1/3, +1/3, -2/3), whatever happened on other problems.

Different roles. Proposer rewards and solver rewards measure different jobs, so they never share a baseline. Proposers are compared with other proposals made from the same source task.

Zero-sum games. In a two-player game the rewards sum to roughly zero whatever the policy does, so a group mean carries no information. Centering against it also turns a structural asymmetry, such as a first-mover advantage, into permanent credit for one seat. SPIRAL instead keeps a separate moving-average baseline of rewards for each game and role, called role-conditioned advantage estimation.

Co-adaptation

The main failure mode is co-adaptation. Proposer and solver drift toward problems that are easy to generate and verify, strangely formatted, or exploitable in ways the checker misses, while getting no better at the tasks anyone cares about. When the checker cannot tell an ambiguous problem from a hard one, a proposer rewarded for 50% solve rates learns to write problems that solvers get right by chance half the time.

Competitive self-play also cycles when nothing anchors it. Policy A beats B, B′ beats A, A′ beats B′, and the population moves around a loop without improving against outside opponents. A judge trained alongside a solver learns to accept the solver's style instead of checking correctness, and the solver learns to produce that style (reward hacking).

Anchors

Self-play needs an external anchor that the moving game cannot change:

  • Fixed evaluations on held-out tasks, measured on a schedule, show whether gains transfer (evaluating agents).
  • Programmatic verification of proposed problems and solver answers keeps the reward grounded (verifiable rewards).
  • Frozen opponents or reference checkpoints reveal cycling, because an improvement that transfers also beats old versions.
  • Proposal diversity metrics catch a proposer collapsing onto one family of problems.

A multi-agent episode yields several traces whose rewards depend on each other. The proposer's reward is unknown until its solvers finish, and a game's reward depends on both players. The training system keeps these traces together until the episode is complete, computes advantages with the right comparison groups, and only then splits them into per-agent samples (rollout generation).