What RL changes: compute, capability, and forgetting
An RL scaling curve plots a policy's reward on held-out tasks of the training distribution against the compute spent training it, and its shape separates two properties of a recipe: the ceiling it approaches and the compute it takes to get there. For agent RL, where one run can cost thousands of GPU-hours, the curve determines which recipe to scale, what a small pilot run can predict, and whether a gain reflects new capability or something the base model could already do.
Sigmoid compute curves
ScaleRL fit reward-versus-compute curves across more than 400,000 GPU-hours of runs and found they follow a saturating sigmoid:
is the starting reward, the asymptote the run approaches, the compute at which half the gain has arrived, and how sharply the curve rises. is the ceiling, while and together set efficiency. The fit is useful because it extrapolates: curves fit on an 8B model up to 50,000 GPU-hours predicted the reward the run reached at 100,000, and fits on the first half of 16,000-GPU-hour ablations predicted the second half.
Worked example
Two recipes start from the same model at . Recipe X fits , , GPU-hours. Recipe Y fits , , , close to the parameters ScaleRL reports for its own 8B recipe.
| GPU-hours | Recipe X | Recipe Y |
|---|---|---|
| 1,000 | 0.41 | 0.35 |
| 2,000 | 0.49 | 0.42 |
| 4,000 | 0.51 | 0.52 |
| 16,000 | 0.52 | 0.60 |
| 100,000 | 0.52 | 0.61 |
An ablation compared at 2,000 GPU-hours picks X by seven points. The curves cross near 4,000, and at 100,000 GPU-hours Y leads by nine. Comparing fitted values from short runs, and not rewards at a fixed step, would have picked Y.
Ceiling and speed
Many design choices move and and leave nearly unchanged. In ScaleRL's leave-one-out ablations, removing any single component from the final recipe, such as its loss aggregation, advantage normalization or curriculum, left within about 0.02 and changed only efficiency. A few choices moved the ceiling. The loss type did, with CISPO and GSPO well above DAPO, and so did computing the LM head in FP32, which raised from 0.52 to 0.61 by shrinking the trainer–inference mismatch that feeds the importance ratio. Larger batches also raised the asymptote while looking worse early.
Stability over long runs
A fitted curve predicts only runs that do not collapse. Zheng et al. stabilized training by keeping staleness and trainer–inference mismatch small, the condition under which the token-level loss approximates the sequence-level objective. Once training on a 30B MoE model was stable, on-policy and off-policy variants reached similar peak performance. In a separate experiment on a smaller MoE, three cold-start checkpoints distilled from different teacher models converged to comparable results.
Long runs add their own problems. ProRL trained a 1.5B model for more than 2,000 steps using KL control, periodic reference resets and a broad task mix. When a ProRL model plateaued after 3,000 training steps, BroRL resumed its improvement by raising the number of rollouts per example from 16 to 512, broadening exploration at each step.
Whether RL adds capability
pass@k at large asks whether a model can solve a task at all, given many tries. Yue et al. found that RLVR-trained models beat their base models at small while the base models caught up and overtook them at large . The arithmetic shows how. A base model that solves a task with probability 0.01 has pass@256 of . If RL drives that task's probability to zero while lifting others, the RL model's pass@1 rises and its pass@256 on that task falls to zero. On this reading, RL mostly sharpens the distribution toward solutions the base model could already sample.
Two results complicate that reading. ProRL reports its prolonged-training models beating the base across a wide range of , including tasks where the base model fails however many attempts it gets. Wen et al. argue that pass@k on math overcounts the base model, because at large a correct final answer can come from a wrong derivation. Their CoT-pass@k counts a sample only when both the reasoning and the answer are correct, and under it RLVR extends the reasoning boundary. The practical check is to report pass@k at several values of (evaluation) and to read samples at the largest one.
Spurious and minimal signals
A reward curve can rise for reasons unrelated to the reward. Shao et al. trained Qwen2.5-Math-7B with random rewards and gained 21.4 points on MATH-500, against 29.1 with ground-truth rewards. The authors attribute the gain to GRPO's clipping bias amplifying a behavior already in the model, reasoning written as code, which rose from 65% to over 90% of responses. The same spurious rewards mostly failed to help Llama3 or OLMo2. A single training example can produce a similar jump on the same model family (cold start). Wu et al. traced part of the Qwen2.5 results to benchmark contamination and found that on a newly generated, leakage-free arithmetic set, only accurate rewards produced steady gains.
These effects can masquerade as learning when the base model has seen the benchmark or already holds the target behavior. Choosing base models and evaluations where they cannot means training at least two model families, evaluating on tasks generated after the base model's data cutoff, and running a random-reward or format-only control. A gain that the control also shows is elicitation, not something the reward taught.
Forgetting
RL fine-tuning forgets less of the base model's other abilities than supervised fine-tuning does at equal accuracy on the new task. RL's Razor found that forgetting tracked the KL divergence from the base model measured on the new task's inputs, , with across its LLM runs. On-policy updates only reweight what the model already samples, so among the policies that solve the new task, RL tends to reach the ones closest in KL to where it started. SFT pulls toward an external distribution that can be arbitrarily far away.
This makes KL to the base model, measured on training prompts, a cheap running proxy for forgetting in a long agent run. When retaining general ability matters, it favors on-policy methods, including on-policy distillation, over SFT on external data for the late stages of post-training.