2From One Agent to Many
Multi-agent reinforcement learning is not simply single-agent learning repeated for several agents.
In this section you will
- Explain why the value of an action depends on what the other agent does
- Describe why a learning partner makes the environment non-stationary
- State the scope of the cooperative setting used throughout this resource
- Name the three capabilities the rest of the resource develops
Two agents now make simultaneous decisions in the same kitchen and receive one shared reward.
Consequences of Multiple Learners
Section titled “Consequences of Multiple Learners”The single-agent interaction loop changes only by adding a second learner.
Notice what is not there: no arrow between the two agents. They never touch each other directly. Everything one agent does reaches the other only by changing the kitchen they share.
Two things, and they are the reason this subject exists.
My outcome depends on your actions. I can choose perfectly and still get a bad result.
If you are learning too, the world I am learning in keeps changing. The thing I am adapting to is itself adapting.
The rest of this section takes those one at a time.
Interdependent Actions
Section titled “Interdependent Actions”Consider two steps in the kitchen, and what the team gets from each.
| Agent 1 does | Agent 2 does | What the team gets |
|---|---|---|
| place pot on stove | bring the ingredient | the pot heats while the ingredient arrives: progress |
| bring the ingredient | bring the ingredient | two agents, one ingredient, cold stove: duplicated effort |
| place pot on stove | place pot on stove | one of them is redundant: wasted step |
Look at the first column. “Bring the ingredient” is an excellent action in row one and a wasted one in row two, and the agent doing it did exactly the same thing both times.
That is the whole difficulty in one observation: an action can no longer be evaluated on its own. In the single-agent kitchen, if the ingredient was not brought, the agent did not bring it. Now, whether bringing it was the right call depends on a decision being made simultaneously by somebody else.
So the object the environment responds to is not any one agent’s action. It is all of them together.
- the joint action: what everybody did this step, in bold
- the action agent i chose, from its own information
- the number of agents
The next section builds the environment around this object properly. For now only the shape matters: rewards and transitions will be functions of , never of alone.
Non-Stationarity During Learning
Section titled “Non-Stationarity During Learning”The second change is subtler and does more damage.
In single-agent reinforcement learning the environment is a fixed thing. It may be random, but its rules do not move. Learn them well enough and they stay learned.
Now put yourself inside agent 1. From where you stand, agent 2 is part of your environment. It is one of the things that determines what happens when you act. And agent 2 is learning.
Yesterday, passing an ingredient to agent 2 worked: it was waiting at the stove, it took the ingredient, the order went out. Today agent 2 has learned to fetch its own ingredients, so it is not at the stove any more. The same action, in the same situation, now fails.
Nothing in the kitchen changed. The stove works the same way. What changed is a partner’s policy, and from agent 1’s point of view that is indistinguishable from the environment’s rules changing underneath it.
This breaks something that single-agent methods quietly rely on: that experience collected earlier still describes the world you are in now. It does not, and the staler the experience, the less it describes. Independent Learning works through the consequences and what can be done about them. Here it is enough to see where the problem comes from.
Knowledge check
Correct.
Not quite.
Its partner learned something different, so the same actions now produce different outcomes.
Yes. This is non-stationarity from the inside. Agent 1 is holding still in an environment that includes a learner, so its returns can drift without anything it controls having changed.
Agent 1 must have a bug, since a fixed policy in a fixed environment gives fixed returns.
That reasoning is right for single-agent RL and wrong here. The environment is not fixed: it contains another learning agent, whose policy is part of what determines agent 1’s outcomes.
The reward function changed.
Possible in principle, but not needed to explain this. The reward function can be completely fixed and returns will still drift, because the reward depends on the joint action and the other half of that joint action is changing.
The discount factor is too low.
The discount factor changes what the agent is optimising, not whether its outcomes drift while it holds still. That drift comes from the partner.
Explanation
The useful habit is to stop treating “the environment” as the physical world and start treating it as everything outside this agent, partners included.
The Cooperative MARL Scope
Section titled “The Cooperative MARL Scope”Multi-agent reinforcement learning is the study of several agents that
- interact in a shared environment,
- make their own decisions, each from its own information,
- and learn from what happens.
That definition says nothing about whether they are on the same side, and in general they need not be. Agents can be strictly opposed, as in a two-player game where one wins exactly what the other loses. They can have partly overlapping interests, which is where most of economics and game theory happens.
This resource is about the cooperative case: every agent receives the same team reward and they are trying to achieve one thing together. That is a genuine restriction, and it is also the setting behind warehouse robotics, collaborative search, autonomous traffic, and the cooperative kitchen used here. It is worth understanding before mixed or competitive cases.
Coordination, Communication, and Adaptation
Section titled “Coordination, Communication, and Adaptation”Everything hard about the two-agent kitchen falls under one of three questions, and each of them is a chapter of this resource.
Coordinate. How should agents choose actions that work well together? Their actions interfere, their reward is shared, and nobody can evaluate a choice in isolation. the Coordinate chapter.
Communicate. What should agents share when they know different things? Each agent sees only part of the kitchen. Talking helps, but a channel has limited bandwidth and messages are not free. the Communicate chapter.
Adapt. Can agents cooperate when the agents around them change? A team that trained together can be very good at working with each other and useless with anybody else. the Adapt chapter.
- select compatible actions under a shared objective
- exchange decision-relevant local information
- cooperate when partner behaviour changes
Multi-Agent Learning Summary
Section titled “Multi-Agent Learning Summary”- Multi-agent reinforcement learning is not single-agent learning run times. Adding one learner changes the problem, not just its size.
- Actions are interdependent: the same action can be right or wrong depending on what a partner chose at the same moment. So the environment responds to the joint action .
- Learning agents make each other’s environments non-stationary. Held fixed, an agent’s returns can still drift, because its partners are moving.
- Agents never influence each other directly. Everything travels through the shared environment.
- This resource covers the cooperative case: one shared reward, one team objective. Competitive and mixed settings exist and are out of scope.