Groups, batches, and packing
A training system arranges the same rollouts in three different ways. A group is a set of rollouts of the same task whose rewards are compared to compute advantages. A batch is the set of samples that together produce one optimizer step. Packing is the physical layout that places variable-length token sequences into dense tensors on accelerators. The three often hold overlapping data, but each answers a different question, and letting one determine another changes the algorithm.
Groups, batches, and packed sequences
Groups as comparison sets
A group holds alternative rollouts of one task under the same start conditions, and it is where a group baseline such as GRPO's is computed. Group size is an algorithm choice: larger groups give a lower-variance baseline and a better chance of seeing both successes and failures on hard tasks, at the cost of more generation per task. Each rollout's advantage has to come from its own group, which rollout generation keeps track of with a group identifier.
Batches as update units
A batch is everything that contributes to one gradient step. It can contain many groups, several environments, and sequences of very different lengths. Its size affects gradient variance, memory, and how far the policy moves between weight broadcasts.
For agents, "batch size 128" leaves the scale ambiguous. It could mean 128 rollouts, 128 trainable branches (one rollout can yield several after compaction or subagents), or a token budget. One long agent rollout can hold more tokens than twenty short ones combined. Run metadata that records groups, rollouts, trainable samples and sampled tokens per step explains why two batches with the same nominal size can differ widely in cost and gradient scale.
A worked example. Assume each rollout yields one trainable branch. With a batch size of 128 and a group size of 16, a step nominally holds eight complete groups. Suppose two of those groups have uniform rewards, all zero or all one. Their advantages are all zero and they contribute no policy-gradient signal. A system can train on the six useful groups (96 samples), or drop the uniform groups and keep collecting until 128 samples with signal arrive. The second choice keeps the step's effective size constant, and a step may then draw on more than eight groups.
Groups also finish out of order, so a batch fills with whichever groups complete first. Which groups those are depends on how generation handles slow rollouts (rollout generation).
Packed rows
Accelerators want dense, fixed-shape tensors. Rollouts arrive ragged: 900 tokens, 14,000 tokens, 3,200 tokens. Padding each to the longest wastes most of the compute. Packing concatenates several samples into one fixed-length row and prevents attention from crossing sample boundaries, typically by resetting position IDs and passing sequence boundaries to a variable-length attention kernel (Krell et al. describe packing without cross-contamination).
Packing leaves each sample's meaning unchanged. Tokens from different traces may sit side by side in memory, but each keeps its own loss mask, advantage and behavior log-probability. Across data-parallel ranks, a packer also balances the work so that no rank waits on another that holds all the long sequences. Attention cost grows faster than linearly in length, so balancing by estimated FLOPs spreads the work more evenly than balancing by token count.
Loss normalization across rows
Normalization must not depend on packing layout. Whatever averaging the loss uses, a token's weight in the update should be the same whether its sample shares a row with others, is split across two rows, or lands in a different micro-batch. Averaging each micro-batch separately breaks this: a token in a row holding 500 trainable tokens then counts four times as much as a token in a row holding 2,000. Dividing by a token count computed over the whole step, across all ranks, makes each token's weight a property of the data.
Mixing environments in one batch
Training often mixes environments so the policy practices several skills and keeps its breadth. Each environment can have its own reward scale, group size and typical sequence length. Pooled without weights, one environment can dominate the update through token volume or reward scale. Explicit sampling weights set each environment's share of the batch, computing advantages within each group keeps one environment's reward offset out of another's baseline, and per-environment logging shows what each contributes.