Cold start and difficulty filtering

The cold-start problem is the case where the initial policy almost never succeeds on the training tasks, so nearly every group of rollouts fails uniformly and produces no learning signal. In agent RL it typically appears on a new environment or harness, where a policy that has not learned the tool format fails every task for the same reason. Difficulty filtering is the matching response on the task side: estimate each task's pass rate and train where outcomes vary.

Why all-fail groups teach nothing

With a group-relative baseline, a group whose rewards are all equal has all-zero advantages (uniform groups), so the policy gradient for that task is zero however much compute the rollouts consumed. For a binary task with success probability pp and group size nn, the chance that a group contains at least one success is 1−(1−p)n1-(1-p)^n. At p=0.01p = 0.01 and n=8n = 8 that is about 8%, so more than nine of every ten groups are generated and then discarded by the trainer.

Pass rate and learning signal are therefore different measurements. A task distribution concentrated on trivial problems shows high reward and almost no gradient, and one concentrated on impossible problems shows zero reward and no gradient.

Advantage mass

Take a binary group of nn rollouts with kk successes and observed success fraction p^=k/n\hat p = k/n. Under mean-centering each success gets 1−p^1-\hat p and each failure −p^-\hat p, so the group's total advantage mass is

∑i∣Ai∣=k(1−p^)+(n−k)p^=2np^(1−p^).\sum_i |A_i| = k(1-\hat p) + (n-k)\hat p = 2n\hat p(1-\hat p).

For independent attempts with true success probability pp, E[p^(1−p^)]=n−1np(1−p)\mathbb{E}[\hat p(1-\hat p)] = \tfrac{n-1}{n}p(1-p), so the expected mass is 2(n−1)p(1−p)2(n-1)p(1-p). With n=8n = 8:

ppExpected massShare of the maximum
0.53.5100%
0.11.2636%
0.010.144%

The mass peaks at p=1/2p = 1/2, where uniform groups are also least likely, which is the basis for the common target of tasks near 50% success. The target is a heuristic. Dividing by the group's standard deviation shifts weight toward both extremes and dividing by its mean shifts it toward hard tasks (normalization choices), and group size, reward shape and per-rollout cost all move the useful band.

Warm-starting the policy

The most direct fix is to make the policy better before RL. Supervised fine-tuning on demonstrations, rejection-sampled successes from a stronger model, or on-policy distillation can move the pass rate on the target distribution from near zero into a range where groups are mixed. For agents, a short SFT warm-start mostly teaches the harness's tool-call format and basic workflow. Those failures are unrelated to task difficulty, and a policy that emits malformed calls earns zero on easy and hard tasks alike, so filtering cannot find a learnable band until the format is fixed.

Making tasks easier

The other lever is the task distribution:

  • Curricula start with easier tasks and raise difficulty as the policy improves.
  • Hints or partial solutions in the prompt are withdrawn over training. A teacher with privileged context is the distillation analogue (teachers and privileged context).
  • Partial-credit rewards break the all-zero tie but add intermediate proxies the policy can exploit (reward hacking).

Difficulty filtering

Offline filtering runs the starting policy several times per task and drops tasks it always or never solves. In the INTELLECT-2 ablations, GRPO on DeepSeek-R1-Distill-Qwen-7B with the unfiltered DeepScaleR math set barely improved reward. Keeping only problems the base model solved in 1 to 4 of 8 attempts made reward rise, and the same 7B model then prefiltered the training data for the INTELLECT-2 run itself.

Online filtering acts on groups after generation. DAPO popularized dynamic sampling: over-sample prompts, discard groups whose accuracy is 0 or 1, and keep sampling until the batch is full of groups with nonzero advantage. INTELLECT-2 used the same rule. This keeps the effective batch size constant, so gradient noise does not drift as more tasks saturate. It restores signal density without saving compute: if a fraction qq of groups is informative, filling a batch of BB informative groups costs about B/qB/q groups of generation. ScaleRL (Khatri et al.) separates two cheaper variants. Dropping zero-variance groups from the batch without resampling gave a higher asymptotic pass rate than keeping them, and permanently removing prompts once their historical pass rate reaches 0.9 (no-positive-resampling) also raised it.

Filtering should not collapse the training distribution onto a narrow band permanently. The frontier moves as the policy improves, and tasks dropped as impossible early may become learnable. Evaluation still needs easy and hard tasks to measure coverage.

How few tasks can suffice

Filtering shrinks the taskset, and a small set of well-chosen tasks can carry a run. Wang et al. trained Qwen2.5-Math-1.5B with RLVR on one math problem and raised MATH500 accuracy from 36.0% to 73.6%, matching training on a 1.2k-problem set that contained it. A control rewarded only for a parseable answer reached 65.0%, so 8.6 points of the gain go beyond format correction. Qwen2.5-Math is a family whose math RL gains are hard to separate from pretraining exposure (RL scaling), so the result shows how little data can unlock behavior the base model already has, not that one task teaches a new skill.

Tracking the informative fraction

A static dataset moves from learnable to useless as the policy improves. Logging the fraction of groups with nonzero advantage alongside reward shows when that is happening. A falling fraction with rising reward means the tasks are saturating, and a fraction stuck near zero with flat reward means the run is cold.