Skip to content
MARL in Cooperative Environments
Edit this page

2.2Communication in Cooperative MARL

6 min read

Communication addresses information asymmetry: cooperative agents may need one another’s private observations or intentions to choose compatible actions.

In this section you will

  • Define information asymmetry in a cooperative team
  • Explain which coordination failures local information alone cannot fix
  • Distinguish communication at execution from centralized training
  • Treat an incoming message as an input to a policy

Back to the kitchen, with the partition up.

Two robot agents working in one kitchen. The left agent is at a chopping board, the right agent at a stove with a pot. A partition stands between them, so neither can see what the other is doing. Agent 1Agent 2neither can see past this

Agent 1 can see the order ticket. Agent 2 can see the stove. Neither can see the other’s half.

Agent 1 knows the next order is soup. Agent 2 knows the current dish is nearly done. Both are maximising the same team reward, and neither knows the thing the other knows.

This is not simply partial observability. Partial Observability established that an agent’s view is a crop of the truth, and that some states are indistinguishable from inside. That is a statement about one agent’s ignorance.

The multi-agent version is sharper, and it is the premise of this chapter:

That distinction changes what a solution can look like. If nobody knows the thing, an agent has to infer it, or hedge, or accept the loss. If a partner knows it, there is a third option: ask, or be told.

Without communication, coordinated behaviour has to come from somewhere. the Coordinate chapter used all of these, mostly without naming them:

  • local observations, react to what you can see,
  • memory, an agent’s own history disambiguates its current view,
  • learned conventions, fixed role assignments that happen to fit together, like agent 1 always fetching,
  • predictability, modelling what a partner tends to do, which works precisely while the partner is stable.

Communication adds a fifth source, and it is the only one that moves information between agents at execution time.

But it is worth setting the bar immediately, because this chapter can very easily degenerate into “more messages are better”.

This distinction deserves its own section, because the Coordinate chapter just spent four sections on CTDE and the two ideas are easy to run together.

Under CTDE, extra information ztz_t is available to the learning algorithm. It shapes gradients during training and is gone at deployment. The deployed policy is a function of local information alone.

Under communication, one agent sends something another agent uses while acting. It exists at deployment, by definition. That is the whole point of it.

Centralized training (CTDE)Communication
When is the information used?during trainingduring execution
Who consumes it?the learning algorithmanother agent’s policy
Does it survive deployment?noyes
Does it need a channel?noyes, with capacity, delay, loss
Does it cost anything at run time?nothingbandwidth, energy, airtime

The two also compose, and most modern systems use both: centralized training to learn what to say, and a channel at execution to say it. the Adapt chapter’s research connection is a good example, the protocol is learned with centralized help and then has to survive on its own.

One more boundary before the formalism. A message does not directly change the environment.

If agent 1 says “soup next”, the stove does not get hotter and no tomato moves. The kitchen is exactly as it was. What changes is what agent 2 knows, and therefore what agent 2 may choose to do next, and only then, through agent 2’s action, does the kitchen change.

Figure 1
environment action⟶the kitchenmessage⟶a partner’s decision\tone{action}{\text{environment action}} \longrightarrow \tone{observe}{\text{the kitchen}} \qquad\qquad \tone{comm}{\text{message}} \longrightarrow \tone{policy}{\text{a partner's decision}}
environment action\text{environment action}
changes the environment through the transition function
message\text{message}
changes information available to another policy
Environmental actions change state; messages change another agent\u2019s decision input.

This is why communication can be modelled as an extra action with an unusual property: other agents observe it, and the state does not depend on it. The next section makes that precise.

Knowledge check

A team trains with a critic that sees every agent's observation. At deployment, each agent acts on its own observation and sends nothing. Is this communication?

Select one answer.

  • Partners suffer information asymmetry: each holds something the other could use. The information is in the system, in the wrong place.
  • Without a channel, coordination relies on local observations, memory, learned conventions and partner predictability.
  • Communication is useful only when it changes a useful decision. Messages the receiver already knew, cannot act on, or would ignore earn nothing.
  • CTDE is not communication. Centralized information is consumed by the learning algorithm and disappears; a message is consumed by another agent’s policy at execution and needs a channel.
  • The radio test: remove the radios from the deployed system. If it still works, it was CTDE.
  • A message does not change the environment. It changes what a partner knows, and so what a partner might do.