Skip to content
MARL in Cooperative Environments
Edit this page

2From One Agent to Many

7 min read

Multi-agent reinforcement learning is not simply single-agent learning repeated for several agents.

In this section you will

  • Explain why the value of an action depends on what the other agent does
  • Describe why a learning partner makes the environment non-stationary
  • State the scope of the cooperative setting used throughout this resource
  • Name the three capabilities the rest of the resource develops
Two robot agents working in one kitchen. The left agent is at a chopping board, the right agent at a stove with a pot. They share one counter. Agent 1Agent 2

Two agents now make simultaneous decisions in the same kitchen and receive one shared reward.

The single-agent interaction loop changes only by adding a second learner.

Two agents each act on the same environment, and each receives its own observation and reward from it. Neither agent is connected directly to the other: one agent influences the other only by changing the environment they share.actionobservation, rewardactionobservation, rewardAgent 1Agent 2Environmentagent 1 reaches agent 2 only through the environment

Notice what is not there: no arrow between the two agents. They never touch each other directly. Everything one agent does reaches the other only by changing the kitchen they share.

Two things, and they are the reason this subject exists.

My outcome depends on your actions. I can choose perfectly and still get a bad result.

If you are learning too, the world I am learning in keeps changing. The thing I am adapting to is itself adapting.

The rest of this section takes those one at a time.

Consider two steps in the kitchen, and what the team gets from each.

Agent 1 doesAgent 2 doesWhat the team gets
place pot on stovebring the ingredientthe pot heats while the ingredient arrives: progress
bring the ingredientbring the ingredienttwo agents, one ingredient, cold stove: duplicated effort
place pot on stoveplace pot on stoveone of them is redundant: wasted step

Look at the first column. “Bring the ingredient” is an excellent action in row one and a wasted one in row two, and the agent doing it did exactly the same thing both times.

That is the whole difficulty in one observation: an action can no longer be evaluated on its own. In the single-agent kitchen, if the ingredient was not brought, the agent did not bring it. Now, whether bringing it was the right call depends on a decision being made simultaneously by somebody else.

So the object the environment responds to is not any one agent’s action. It is all of them together.

Figure 1
at=(at1, …, atn)\tone{action}{\jointact} = \bigl(\tone{policy}{\act{1}},\ \dots,\ \tone{policy}{\act{\nag}}\bigr)
at\jointact
the joint action: what everybody did this step, in bold
ati\act{i}
the action agent i chose, from its own information
n\nag
the number of agents
Actions arrive at the environment as a set, and the environment responds to the set.

The next section builds the environment around this object properly. For now only the shape matters: rewards and transitions will be functions of at\jointact, never of ati\act{i} alone.

The second change is subtler and does more damage.

In single-agent reinforcement learning the environment is a fixed thing. It may be random, but its rules do not move. Learn them well enough and they stay learned.

Now put yourself inside agent 1. From where you stand, agent 2 is part of your environment. It is one of the things that determines what happens when you act. And agent 2 is learning.

Yesterday, passing an ingredient to agent 2 worked: it was waiting at the stove, it took the ingredient, the order went out. Today agent 2 has learned to fetch its own ingredients, so it is not at the stove any more. The same action, in the same situation, now fails.

Nothing in the kitchen changed. The stove works the same way. What changed is a partner’s policy, and from agent 1’s point of view that is indistinguishable from the environment’s rules changing underneath it.

This breaks something that single-agent methods quietly rely on: that experience collected earlier still describes the world you are in now. It does not, and the staler the experience, the less it describes. Independent Learning works through the consequences and what can be done about them. Here it is enough to see where the problem comes from.

Knowledge check

Agent 1's policy is unchanged and the kitchen's physics are unchanged, yet the returns it collects have got worse over the last hundred episodes. What is the most likely explanation?

Select one answer.

Multi-agent reinforcement learning is the study of several agents that

  • interact in a shared environment,
  • make their own decisions, each from its own information,
  • and learn from what happens.

That definition says nothing about whether they are on the same side, and in general they need not be. Agents can be strictly opposed, as in a two-player game where one wins exactly what the other loses. They can have partly overlapping interests, which is where most of economics and game theory happens.

This resource is about the cooperative case: every agent receives the same team reward and they are trying to achieve one thing together. That is a genuine restriction, and it is also the setting behind warehouse robotics, collaborative search, autonomous traffic, and the cooperative kitchen used here. It is worth understanding before mixed or competitive cases.

Coordination, Communication, and Adaptation

Section titled “Coordination, Communication, and Adaptation”

Everything hard about the two-agent kitchen falls under one of three questions, and each of them is a chapter of this resource.

Coordinate. How should agents choose actions that work well together? Their actions interfere, their reward is shared, and nobody can evaluate a choice in isolation. the Coordinate chapter.

Communicate. What should agents share when they know different things? Each agent sees only part of the kitchen. Talking helps, but a channel has limited bandwidth and messages are not free. the Communicate chapter.

Adapt. Can agents cooperate when the agents around them change? A team that trained together can be very good at working with each other and useless with anybody else. the Adapt chapter.

Figure 2
Coordinate  ⟶  Communicate  ⟶  Adapt\cbox{action}{\text{Coordinate}} \;\longrightarrow\; \cbox{comm}{\text{Communicate}} \;\longrightarrow\; \cbox{policy}{\text{Adapt}}
Coordinate\text{Coordinate}
select compatible actions under a shared objective
Communicate\text{Communicate}
exchange decision-relevant local information
Adapt\text{Adapt}
cooperate when partner behaviour changes
The course progression from compatible actions to information exchange and partner generalization.
  • Multi-agent reinforcement learning is not single-agent learning run nn times. Adding one learner changes the problem, not just its size.
  • Actions are interdependent: the same action can be right or wrong depending on what a partner chose at the same moment. So the environment responds to the joint action at\jointact.
  • Learning agents make each other’s environments non-stationary. Held fixed, an agent’s returns can still drift, because its partners are moving.
  • Agents never influence each other directly. Everything travels through the shared environment.
  • This resource covers the cooperative case: one shared reward, one team objective. Competitive and mixed settings exist and are out of scope.