Skip to content
MARL in Cooperative Environments
Edit this page

1Reinforcement Learning

8 min read

Reinforcement learning concerns an agent that improves its decisions through interaction rather than labelled examples.

In this section you will

  • Describe the agent and environment loop, one step at a time
  • Distinguish a policy from a single action
  • Compute a discounted return from a sequence of rewards
  • State the objective a single agent optimizes
  • Map a one-agent kitchen onto the loop
One robot agent alone in a kitchen, standing at a counter with a chopping board on one side and a stove with a pot on the other. It must work both stations itself. Agent

One agent, one kitchen, one order to serve. Everything in this section is a question about what this agent should do next.

Reinforcement learning is what you do when nobody can tell you the right answer, but somebody can tell you how well things went.

The setup has two parts and nothing else. The agent is the thing that decides. The environment is everything else, in our case the kitchen, the ingredients, the stove and the order. The agent takes an action, the environment changes and reports back, and the agent decides again. Learning happens in that loop and nowhere else.

Two features make this a distinct kind of problem.

The decisions are sequential, and they interfere with each other. Turning on the stove is not good or bad by itself. It is good if you are about to cook and bad if you have not collected the ingredients, and you will not find out which for several steps. Each choice changes the situation the next choice is made in.

The feedback evaluates rather than instructs. This is the sharpest break from supervised learning, and it is worth stating in Sutton and Barto’s terms: reinforcement learning uses evaluative feedback, supervised learning uses instructive feedback. A label says “the answer was 7”. A reward says “that went badly”, without saying what would have gone well.

ParadigmWhat the data gives youWhat you learn
Supervisedinputs paired with correct outputsto reproduce the labels
Unsupervisedinputs onlystructure in the inputs
Self-supervisedinputs, with part of each hiddento predict the hidden part
Reinforcementa scalar score for what you actually didto act so the score is high

The last row differs from the others in two ways that matter. You only ever learn about the action you took, so you have to try things to find out about them, nobody hands you a dataset of alternatives. And the data you learn from is generated by your own behaviour, so it changes as you improve.

Formally, the interaction is a repeating cycle of four things.

A single agent acts on the environment, and the environment returns an observation and a reward. The two form a closed loop.actionobservation, rewardAgentEnvironment

At each step tt: the agent receives the state st\st, describing the situation; it takes an action ata_t; the environment returns a reward rt\rew, a single number scoring what happened; and the environment moves to the next state st+1s_{t+1}.

Figure 1
st  ⟶  at  ⟶  rt, st+1\tone{observe}{\st} \;\longrightarrow\; \tone{action}{a_t} \;\longrightarrow\; \tone{reward}{\rew},\ \tone{observe}{s_{t+1}}
st\st
the state at step t: what the situation is now
ata_t
the action the agent takes in response
rt\rew
the reward: one number scoring what just happened
st+1s_{t+1}
the next state, which is where the following step begins
One step. Repeat this and you have everything reinforcement learning studies.

Run that cycle from a starting situation until the task ends, the order is served, or the kitchen runs out of time, and you have one episode. Learning happens across many episodes.

The agent’s behaviour is entirely captured by one object: its policy, the rule that turns a situation into an action. When you say a learner has improved, you mean its policy changed.

Figure 2
at∼π(at∣st)\tone{action}{a_t} \sim \tone{policy}{\pi}\bigl(a_t \given \tone{observe}{\st}\bigr)
π\pi
the policy: the agent’s decision rule
at∼a_t \sim
the action is drawn from the distribution the policy defines
st\st
the situation the policy is given
The policy is the agent. Given this situation, what do I do?

Two flavours, and the distinction returns later. A deterministic policy picks one action for each state: at=π(st)a_t = \pi(\st). A stochastic policy defines a distribution over actions and samples from it, as written above. Stochastic policies matter because randomness is how an agent tries things it has not tried, and, in the Coordinate chapter, because sometimes a team is genuinely better off flipping a coin than committing.

An agent that maximises the reward available right now is not doing the task. The kitchen agent could stand at the stove collecting whatever the stove pays and never serve anything. What matters is the whole episode.

Figure 3
Gt=∑k=0T−tγk⏟how much later steps count  rt+k\tone{reward}{G_t} = \sum_{k=0}^{T-t} \ubrace{observe}{\gamma^{k}}{how much later steps count} \; \tone{reward}{r_{t+k}}
GtG_t
the return from step t onward: the total the agent still stands to collect
γ\gamma
the discount factor, between 0 and 1
rt+kr_{t+k}
the reward k steps into the future
TT
the last step of the episode
Not this step's reward. Everything from here to the end of the episode, added up.

The discount factor γ\gamma says how much a reward that arrives later is worth now. At γ=0\gamma = 0 the agent is completely short-sighted and only the current reward exists. As γ\gamma approaches 1 a reward ten steps away counts almost as much as one right now, and the agent becomes willing to invest: collect ingredients that pay nothing in themselves, because they lead to a served order.

Because both the environment and the policy can be random, no single episode is the thing to optimise. The objective is the average return the policy achieves.

Figure 4
J(π)=Eτ∼π[ G(τ) ]\tone{policy}{J(\pi)} = \E_{\tau \sim \tone{policy}{\pi}}\bigl[\, \tone{reward}{G(\tau)} \,\bigr]
J(π)J(\pi)
the objective: how good the policy is
τ\tau
a trajectory: one whole episode of states, actions and rewards
Eτ∼π\E_{\tau \sim \pi}
average over the episodes that policy produces
G(τ)G(\tau)
the return of that episode
A policy is good if the episodes it produces tend to go well.

Everything else in reinforcement learning, value functions, Q-learning, policy gradients, is machinery for finding a π\pi that makes J(π)J(\pi) large. This resource introduces those pieces only where a multi-agent idea actually needs them.

The Kitchen as a Reinforcement-Learning Problem

Section titled “The Kitchen as a Reinforcement-Learning Problem”

Four abstract objects, one concrete kitchen. This is the whole of single-agent reinforcement learning applied to our example.

PieceIn the kitchen
State st\stwhere the agent is, what it holds, which ingredients are prepared, how far the order has got
Action ata_tmove, pick up, put down, chop, place on stove, serve
Reward rt\rewa payment when an order is served
Policy π\pithe agent’s habit of what to do in each situation
Episodeone service, from an empty kitchen to a served order or the end of time
Return GtG_thow much of that service is still to come, discounted

Given all of this, a good policy in the single-agent kitchen is a sequence worked out by one decider: collect, chop, heat, cook, plate, serve. The agent controls every part of it. If something is not done, it is because this agent did not do it.

Knowledge check

The kitchen agent chops a tomato and receives a reward of 0. What has it learned?

Select one answer.

  • Reinforcement learning is learning from interaction: act, observe the consequence, act again.
  • Its feedback is evaluative, not instructive. A reward scores what you did; it does not reveal what you should have done.
  • One step is st→at→rt,st+1\st \rightarrow a_t \rightarrow \rew, s_{t+1}. An episode is a run of those steps.
  • A policy π(at∣st)\pi(a_t \given \st) is the agent. Improving means changing it.
  • The objective is the expected discounted return J(π)J(\pi), not the immediate reward. The discount factor γ\gamma sets how far ahead the agent cares.
  • If the correct actions are already labelled, use supervised learning instead.

Reinforcement Learning: An Introduction, Richard S. Sutton and Andrew G. Barto (2nd edition, MIT Press, 2018). The standard textbook, and the reference this field is written against. The instructive-versus-evaluative distinction above is theirs. Freely available from the authors.

Grokking Deep Reinforcement Learning, Miguel Morales (Manning, 2020). The same ideas through visual explanation and worked code. Often the better first book if Sutton and Barto feels steep.

Reinforcement Learning: A Comprehensive Overview, Kevin P. Murphy (2025, arXiv:2412.05265). A survey rather than a textbook: compact, current, and organised by method. The fastest way to see how any one algorithm relates to the rest. Free.

OpenAI’s Spinning Up in Deep RL is a free set of introductory essays with readable reference implementations of the main policy-gradient algorithms.

For the multi-agent material that starts next, the standard reference is Multi-Agent Reinforcement Learning: Foundations and Modern Approaches by Stefano V. Albrecht, Filippos Christianos and Lukas Schäfer (MIT Press, 2024), free online at marl-book.com. The Resources page maps each of our sections onto it.


One agent, in full control, with a score to maximise. Now hold that picture and change one thing.

What changes when another learning agent enters the same kitchen?