Weight broadcast and policy versions

Weight broadcast is the step in online RL that moves freshly trained parameters from the trainer to the inference engines and makes them live, and a policy version is the label, usually the optimizer step count, that names which parameters are live. Together they determine how stale the sampled data is and whether every token in a trace can be attributed to the weights that produced it.

A broadcast has two parts with different costs. Transfer moves the bytes into inference memory and can run while the engine keeps serving. Activation switches the engine to the new weights at one instant for every request.

Transports

The transfer path depends on scale and topology.

  • Checkpoint files. The trainer writes weights to shared storage and inference loads them. Simple and robust, and slow for large models.
  • Collective communication. Trainer and inference ranks join one communication group, and the trainer broadcasts tensors over NCCL. Fast, but the process group is static, which complicates elastic or fault-tolerant deployments.
  • One-sided RDMA reads. Each inference rank pulls the tensors it needs directly from trainer GPU memory, with libraries such as NIXL, and no rank waits on a collective.

Worked example

A 32B-parameter policy is 64 GB in bf16. Assume, for illustration, shared storage that writes and reads at 2 GB/s. A filesystem broadcast then takes about a minute, 32 seconds to write and 32 to read. On a five-minute step, inference then samples from the previous version for a fifth of every step, which adds lag to every rollout in flight.

Prime Intellect's NIXL weight-transfer write-up measured about 45 GB/s per 400 Gbit/s NIC for RDMA reads. A node with eight such NICs would pull the same 64 GB in about 0.2 seconds at that rate. At that speed activation dominates: in the same write-up, pausing a data- and expert-parallel vLLM deployment took another 3 to 6 seconds, because a rank cannot stop while its peers may still enter an expert-parallel collective, so the ranks have to agree on when to stop.

Layout conversion

Training shards parameters and optimizer state with FSDP and splits experts with expert parallelism. Inference prefers tensor- or expert-parallel replicas with fused, packed, or transposed kernel layouts, a different naming scheme, and possibly FP8 weights. A broadcast gathers shards from one arrangement, converts names and layouts, casts or quantizes, and repartitions into the other, layer by layer to bound peak memory.

The conversion splits into two kinds of operation. Views such as slicing and reshaping only change offsets and strides, so the receiver can read the right bytes directly from trainer memory. Operations that allocate new storage, such as dtype casts, quantization, and packing, have to run on the inference GPU after the bytes arrive. The NIXL write-up records the whole chain once, pulls through the views, and replays the rest on arrival, overlapping the transfer of one group of tensors with the replay of the previous one.

Quantizing on the way to inference means the served model is no longer the trained one at the same version. That difference is part of the trainer–inference mismatch.

Atomic activation

If a request runs while half the layers hold version 18 and half hold version 17, it samples from a model that neither checkpoint describes, and no recorded version identifies it. Two patterns make the switch atomic.

  • Double buffering. One copy of the weights serves while a second receives the update, and serving flips to the new copy when it is complete. No pause, at the cost of twice the weight memory.
  • Pause, swap, resume. Stop stepping the engine, load the new weights, then continue. No extra memory, and generation halts briefly.

Small adapters change the tradeoff. A LoRA adapter is small enough to load beside the serving copy and swap in between engine steps without a pause, which is double buffering at adapter size.

What happens to requests already in flight is a separate choice between draining them on the old weights and resuming them on the new ones, compared in asynchronous RL. The fate of their cached keys and values is covered in KV cache and prefix reuse.

Recording versions

A version label only helps if it travels with the data. Under pinned rollouts one number per rollout suffices. Under in-flight updates the record is a span:

rollout:          code-task-1842
group dispatched: v17
policy span:      v17 → v19
sampled_logprobs: [-0.12, -2.31, -0.04, ...]

The span gives the rollout's age range, and the per-token log-probabilities give the behavior probabilities whichever version produced each token. A complete identity also names the base checkpoint, tokenizer and renderer versions, sampling settings, and serving precision.

Cadence

Broadcasting after every optimizer step keeps sampling closest to training and is the simplest choice when the transfer is short relative to a step. Broadcasting less often saves transfer time and pauses but lets more data arrive from older policies, which importance-ratio and mismatch statistics measure more directly than version distance.