1Reinforcement Learning
Reinforcement learning concerns an agent that improves its decisions through interaction rather than labelled examples.
In this section you will
- Describe the agent and environment loop, one step at a time
- Distinguish a policy from a single action
- Compute a discounted return from a sequence of rewards
- State the objective a single agent optimizes
- Map a one-agent kitchen onto the loop
One agent, one kitchen, one order to serve. Everything in this section is a question about what this agent should do next.
Learning from Interaction
Section titled “Learning from Interaction”Reinforcement learning is what you do when nobody can tell you the right answer, but somebody can tell you how well things went.
The setup has two parts and nothing else. The agent is the thing that decides. The environment is everything else, in our case the kitchen, the ingredients, the stove and the order. The agent takes an action, the environment changes and reports back, and the agent decides again. Learning happens in that loop and nowhere else.
Two features make this a distinct kind of problem.
The decisions are sequential, and they interfere with each other. Turning on the stove is not good or bad by itself. It is good if you are about to cook and bad if you have not collected the ingredients, and you will not find out which for several steps. Each choice changes the situation the next choice is made in.
The feedback evaluates rather than instructs. This is the sharpest break from supervised learning, and it is worth stating in Sutton and Barto’s terms: reinforcement learning uses evaluative feedback, supervised learning uses instructive feedback. A label says “the answer was 7”. A reward says “that went badly”, without saying what would have gone well.
| Paradigm | What the data gives you | What you learn |
|---|---|---|
| Supervised | inputs paired with correct outputs | to reproduce the labels |
| Unsupervised | inputs only | structure in the inputs |
| Self-supervised | inputs, with part of each hidden | to predict the hidden part |
| Reinforcement | a scalar score for what you actually did | to act so the score is high |
The last row differs from the others in two ways that matter. You only ever learn about the action you took, so you have to try things to find out about them, nobody hands you a dataset of alternatives. And the data you learn from is generated by your own behaviour, so it changes as you improve.
The Reinforcement Learning Loop
Section titled “The Reinforcement Learning Loop”Formally, the interaction is a repeating cycle of four things.
At each step : the agent receives the state , describing the situation; it takes an action ; the environment returns a reward , a single number scoring what happened; and the environment moves to the next state .
- the state at step t: what the situation is now
- the action the agent takes in response
- the reward: one number scoring what just happened
- the next state, which is where the following step begins
Run that cycle from a starting situation until the task ends, the order is served, or the kitchen runs out of time, and you have one episode. Learning happens across many episodes.
Policies as Decision Rules
Section titled “Policies as Decision Rules”The agent’s behaviour is entirely captured by one object: its policy, the rule that turns a situation into an action. When you say a learner has improved, you mean its policy changed.
- the policy: the agent’s decision rule
- the action is drawn from the distribution the policy defines
- the situation the policy is given
Two flavours, and the distinction returns later. A deterministic policy picks one action for each state: . A stochastic policy defines a distribution over actions and samples from it, as written above. Stochastic policies matter because randomness is how an agent tries things it has not tried, and, in the Coordinate chapter, because sometimes a team is genuinely better off flipping a coin than committing.
Return and the Learning Objective
Section titled “Return and the Learning Objective”An agent that maximises the reward available right now is not doing the task. The kitchen agent could stand at the stove collecting whatever the stove pays and never serve anything. What matters is the whole episode.
- the return from step t onward: the total the agent still stands to collect
- the discount factor, between 0 and 1
- the reward k steps into the future
- the last step of the episode
The discount factor says how much a reward that arrives later is worth now. At the agent is completely short-sighted and only the current reward exists. As approaches 1 a reward ten steps away counts almost as much as one right now, and the agent becomes willing to invest: collect ingredients that pay nothing in themselves, because they lead to a served order.
Because both the environment and the policy can be random, no single episode is the thing to optimise. The objective is the average return the policy achieves.
- the objective: how good the policy is
- a trajectory: one whole episode of states, actions and rewards
- average over the episodes that policy produces
- the return of that episode
Everything else in reinforcement learning, value functions, Q-learning, policy gradients, is machinery for finding a that makes large. This resource introduces those pieces only where a multi-agent idea actually needs them.
The Kitchen as a Reinforcement-Learning Problem
Section titled “The Kitchen as a Reinforcement-Learning Problem”Four abstract objects, one concrete kitchen. This is the whole of single-agent reinforcement learning applied to our example.
| Piece | In the kitchen |
|---|---|
| State | where the agent is, what it holds, which ingredients are prepared, how far the order has got |
| Action | move, pick up, put down, chop, place on stove, serve |
| Reward | a payment when an order is served |
| Policy | the agent’s habit of what to do in each situation |
| Episode | one service, from an empty kitchen to a served order or the end of time |
| Return | how much of that service is still to come, discounted |
Given all of this, a good policy in the single-agent kitchen is a sequence worked out by one decider: collect, chop, heat, cook, plate, serve. The agent controls every part of it. If something is not done, it is because this agent did not do it.
Knowledge check
Correct.
Not quite.
That chopping led to no immediate payment. Not that chopping was the wrong action.
Right. The reward is evaluative, and it is about this step only. Chopping may be essential to a served order several steps later, which is exactly what the discounted return is for.
That chopping was wrong, since a correct action would have been rewarded.
That reads the reward as a label. A zero reward says this step paid nothing, not that the action was a mistake. Most steps in a long task pay nothing.
That it should have taken a different action, though not which one.
Closer, but still too strong. The agent has learned nothing about the other actions, and nothing yet about whether this one was bad. It only knows what this step paid.
Nothing, because zero rewards carry no information.
A zero is informative, it rules out an immediate payment here. Over many episodes the pattern of where rewards do and do not arrive is what the agent learns from.
Explanation
Confusing “this paid nothing” with “this was wrong” is the most common early mistake in reinforcement learning, and it gets worse in teams.
Reinforcement Learning Summary
Section titled “Reinforcement Learning Summary”- Reinforcement learning is learning from interaction: act, observe the consequence, act again.
- Its feedback is evaluative, not instructive. A reward scores what you did; it does not reveal what you should have done.
- One step is . An episode is a run of those steps.
- A policy is the agent. Improving means changing it.
- The objective is the expected discounted return , not the immediate reward. The discount factor sets how far ahead the agent cares.
- If the correct actions are already labelled, use supervised learning instead.
Further reading
Section titled “Further reading”Reinforcement Learning: An Introduction, Richard S. Sutton and Andrew G. Barto (2nd edition, MIT Press, 2018). The standard textbook, and the reference this field is written against. The instructive-versus-evaluative distinction above is theirs. Freely available from the authors.
Grokking Deep Reinforcement Learning, Miguel Morales (Manning, 2020). The same ideas through visual explanation and worked code. Often the better first book if Sutton and Barto feels steep.
Reinforcement Learning: A Comprehensive Overview, Kevin P. Murphy (2025, arXiv:2412.05265). A survey rather than a textbook: compact, current, and organised by method. The fastest way to see how any one algorithm relates to the rest. Free.
OpenAI’s Spinning Up in Deep RL is a free set of introductory essays with readable reference implementations of the main policy-gradient algorithms.
For the multi-agent material that starts next, the standard reference is Multi-Agent Reinforcement Learning: Foundations and Modern Approaches by Stefano V. Albrecht, Filippos Christianos and Lukas Schäfer (MIT Press, 2024), free online at marl-book.com. The Resources page maps each of our sections onto it.
One agent, in full control, with a score to maximise. Now hold that picture and change one thing.
What changes when another learning agent enters the same kitchen?