Introduction

Reinforcement learning (RL) for LLM agents trains a language model by letting it attempt tasks, scoring each attempt, and updating the weights so that higher-scoring behavior becomes more likely. In this book, RL means online policy-gradient RL: the model generates its own training data, a scorer assigns each attempt (a rollout) a reward, and a gradient step raises the probability of the tokens in better-than-expected rollouts.

The training signal is a reward instead of a target text to imitate. Supervised fine-tuning needs a demonstration of the right answer, while RL needs only a way to recognize a good outcome: a test suite, a checker for a math answer, a program that inspects a database, or a judge model with a rubric. For many agent tasks, writing the check is far cheaper than writing the demonstration.

The axes of the design space

Names such as PPO, GRPO, RLHF and RLVR label particular combinations of a few choices. The book is organized around those choices.

  • Rewards. What counts as success and how a scorer measures it. Reward design is usually the first and largest decision in an agent RL project (rewards).
  • Advantages. How a reward becomes a weight on each rollout or token, measured against a baseline for the task (advantages).
  • Loss functions. How advantages, importance ratios and trust regions combine into the quantity the trainer minimizes (the policy-gradient loss).
  • On-policy or off-policy. Whether the data came from the policy being updated. Asynchronous training is a little off-policy: rollouts lag the trainer by a bounded number of steps, and importance weighting corrects the gap (on-policy and off-policy RL).
  • Student or teacher. Whether the signal comes from the model's own rewards or from another model. A teacher can score the student's own rollouts token by token, which is on-policy distillation, or supply rollouts it wrote itself, which is supervised fine-tuning on off-policy tokens.

The agent and its environment

An agent acts in an environment. The environment presents a task and an initial observation, the agent acts, the environment changes, and the result becomes the next observation, until the task ends (agents and environments).

For an LLM agent, the model generates tokens and a harness turns them into an interaction: it builds each model request, exposes tools, parses output into actions, and returns results. One model generation is a turn. A rollout runs from task to termination, and its stored record is a trace. A question-answer task ends after one turn, while a coding agent may run for dozens of turns with tool calls, file edits and compaction.

The learning loop

Online RL alternates generation and training. Generation runs the current model on tasks and scores the resulting rollouts. Training uses those scored rollouts to update the model, and the updated model generates the next batch.

The learning loop

In this synchronous example, the model is fixed during rollout generation. Training updates it after completed rollouts have been scored and assembled into a batch.

One iteration has six steps:

  1. Sample tasks, usually with several rollouts per task.
  2. Generate rollouts with the current policy inside an environment that executes the agent's actions.
  3. Score each completed rollout with a reward.
  4. Compute advantages, which say how much better or worse each rollout did than expected for its task.
  5. Update the weights with a policy-gradient loss that raises the probability of tokens in above-average rollouts and lowers it in below-average ones.
  6. Publish the new weights to the inference servers that generate the next rollouts.

In a synchronous loop these steps run in order. Production systems overlap them, so generation continues while the trainer works and rollouts come from weights one or more steps old (asynchronous RL).

Worked example

Take a coding task. The agent receives a small repository and an instruction: test_parse_negative fails, fix it. The harness offers a shell tool. One rollout might look like this:

task    tests/test_parse.py::test_parse_negative fails. Fix it.
turn 1  model → shell: pytest -x tests/test_parse.py
        env   ← FAILED test_parse_negative: parse("-3") returned 3
turn 2  model → shell: sed -n 1,40p src/parse.py
        env   ← def parse(s): ... return int(s.lstrip("-")) ...
turn 3  model → edit src/parse.py (keep the sign)
        env   ← ok
turn 4  model → shell: pytest
        env   ← 12 passed
turn 5  model → "Fixed sign handling in parse()."

After the agent stops, the environment runs a private test suite on the final repository. All private tests pass, so the reward is 1.

Sampling the same task four times gives rewards 1,0,0,11, 0, 0, 1: two rollouts fix the bug, one edits the test file and is caught by the private suite, and one runs out of turns. Subtracting the group mean of 0.50.5 gives advantages of +0.5,−0.5,−0.5,+0.5+0.5, -0.5, -0.5, +0.5, so the update raises the probability of the tokens in the two fixes and lowers it in the other two. The prompt and tool output stay in context with no gradient term of their own, because the model did not generate them (the policy-gradient loss).

What makes agents harder

A single-turn task is the horizon-one case of the same loop. Agents stretch each part of it.

  • One final reward covers many decisions, such as which file to open, how to recover from a failed command and what patch to write, so the signal for any one decision is noisy (credit assignment).
  • The environment offers more ways to score well without doing the task. The agent can edit tests, exploit leftover state or satisfy a weak check, and optimization finds these paths (reward hacking).
  • Rollouts are long and uneven. Tool execution, long contexts and a heavy tail of slow rollouts dominate cost, and the training system has to keep inference busy while environments work.

The five parts

The interaction covers what is trained and what it acts on, starting from agents and environments and moving through harnesses, RL environments and building environments at scale to rollouts and traces. Rewards and evaluation covers how completed work becomes a score and how to check that the score means something, from rewards and verifiable rewards to reward hacking and evaluation.

Policy optimization turns scores into weight updates. It runs from advantages and GRPO through importance sampling and KL penalties, and ends with what RL changes in a model: how returns scale with compute, which capabilities it adds, and what it makes the model forget. Training systems covers producing experience while the model changes, including rollout generation, token-level rendering and asynchronous RL. Beyond RL places RL among other post-training methods: learning from a teacher's tokens, on-policy distillation and teachers with privileged context.

Each page stands alone and links its prerequisites at first mention, and the glossary points each term to the page that defines it.