Agents and environments
In reinforcement learning, the agent is the decision-maker and the environment is everything the agent acts on. For an LLM agent the split determines which tokens the policy is trained on, which processes are part of the task, and what can be swapped without changing the task.
Draw a boundary around the agent. Observations enter and actions leave. The model sits inside the boundary together with the harness that builds its requests and turns its tokens into actions. The task, its external state, and the processes that respond to the agent sit outside: a repository, a shell, a browser, a database, and the program that scores the result.
Action and observation
State, observation, and transition
The environment has a state: the information that determines what can happen next. An observation is the part of the state that reaches the next model request. An action causes a transition to a new state, and the environment returns a new observation:
A reward may arrive at each step, though agent tasks usually give a single reward after the rollout ends (rewards).
Worked example
One step of a coding task, in which the agent has been asked to fix a failing test, maps onto these terms as follows.
| Element | In the coding task |
|---|---|
| State | 38 files in the repository, installed packages, running processes, and a private test suite |
| Observation | the task text and the output of commands run so far, 3 files read |
| Action | pytest -x tests/test_parse.py |
| Transition | tests run, processes start and exit, the repository is unchanged |
| Observation | FAILED test_parse_negative: parse("-3") returned 3 |
| Reward | 0 at this step; 1 at the end if the private tests pass |
The agent has seen 3 of 38 files and none of the private tests. Reading a file or running the tests changes what the agent knows without changing the repository, while an edit changes the repository.
MDPs and POMDPs
A Markov decision process (MDP) is the standard model of this interaction. It consists of a set of states , a set of actions , a transition distribution , a reward function , and a distribution over initial states. "Markov" means the next state depends only on the current state and action. Sutton and Barto is the standard reference.
Agents rarely see the full state, so the better model is a partially observable MDP (POMDP), which adds an observation distribution . The agent acts on what it has observed so far. Searching a repository, running a test and asking a subagent are information-gathering actions, because an agent cannot use a file it has not read or an error that was not returned.
An LLM agent carries its history of observations and actions in the context window and conditions on all of it, so the context serves as its working estimate of the hidden state. When the history outgrows the window, the harness selects what to keep (compaction).
Token-level and turn-level views
At the token level, the state is the prompt plus the tokens generated so far, the action is the next token, and the transition appends that token. These transitions are deterministic, and they fully describe a single-turn math problem, where the environment enters only through the prompt and the reward.
At the turn level, the action is one model generation, such as a tool call, a patch or a final message. The transition is whatever the environment does with it, which can be stochastic and slow: a test may be flaky, a web page may change, a subprocess may time out.
Agent RL uses both views. The environment only sees complete turns, while the loss is computed over the tokens that made them up, so many token-level transitions sit inside each turn-level transition (policies).
Termination
A rollout ends when the agent submits an answer, when the environment detects success or an unrecoverable state, or when a budget on turns, tokens or time runs out. A budget ending is a truncation: the task was not finished, so the final state says little about whether the agent would have succeeded, and it should be recorded separately from a scored ending (rollouts and traces).
Most LLM RL treats each rollout as finite and uses no discount (), so a reward at the end counts fully for every earlier action (advantages).
Where the boundary sits in software
The agent/environment split describes roles, not where processes run. A sandbox may run on the same machine as the harness or on a remote service, and a judge may run after the interaction ends. These components belong to the environment when they determine what the agent observes, what its actions do, or how the task ends.
Environment libraries often ship a default harness with each task so that it runs without extra configuration. The harness still belongs to the agent: replacing it changes the agent's behavior on an unchanged task, which is why comparisons hold it fixed (harnesses). The same boundary lets one environment evaluate a frontier API model and then train an open model on unchanged tasks.