Skip to content
MARL in Cooperative Environments
Edit this page

2.6Communication Policies

5 min read

A communication policy determines what an agent sends, whether it sends at a particular time, and which agents receive the message.

In this section you will

  • Select what to send from a local observation
  • Decide when a message is worth sending
  • Choose which agents should receive it
  • Price communication inside the shared reward

An agent’s message is produced by a policy, exactly like its environment action, and from the same input.

Figure 1
mti∼πim(⋅∣hti)\tone{comm}{m^\ag_t} \sim \tone{policy}{\pol{\ag}^{m}}\bigl(\cdot \given \tone{observe}{h^\ag_t}\bigr)
πim\pol{\ag}^{m}
the communication policy: this agent’s rule for what to say
htih^\ag_t
its own history, since what is worth saying often depends on what happened earlier
Saying something is a decision, made locally, from the same information as any other.

Two things follow from that shape. The communication policy is local: an agent can only report what it knows. And it is learnable: nothing about this is hand-written unless you choose to hand-write it, which is Learning Communication Protocols.

Note the input is htih^\ag_t rather than oti\obs{\ag}. Some of the most useful messages are about change, not state. “The order just changed” needs a memory of the previous order to say at all.

Because the message space contains ∅\varnothing, one available message is to send nothing. So the agent faces two questions, and they are genuinely different:

What should I say?

Do I need to say anything at all?

The second is easy to overlook and does most of the work in a constrained system. Most steps in most tasks are unremarkable. If the situation has not changed and the partner’s plan is still right, a message repeating that is pure cost.

That has a sharp edge, and Communication Constraints named it: a protocol where silence means something breaks under message loss. Making silence meaningful buys efficiency and pays for it in fragility.

Three arrangements, in increasing specificity.

Broadcast. Everyone receives it. Simple, and it scales badly: bandwidth grows with the number of listeners, and a message useful to one partner is noise to the rest.

Range-limited. Only nearby agents receive it, as described in Communication Constraints. Who hears you is determined by the state rather than chosen.

Targeted. The sender picks a recipient. More expressive, and it enlarges the message space, since the agent now chooses both content and address.

The final piece, and the one that makes the design questions bite.

If sending is free, the answer to “should I say anything?” is always yes: another message can only add information, so a team should talk constantly. That makes the interesting question disappear.

Charging for it brings the question back.

Figure 2
rt′=rt−λ ct⏟what talking cost\tone{reward}{r'_t} = \tone{reward}{r_t} - \ubrace{comm}{\lambda\, c_t}{what talking cost}
rt′r'_t
the reward the team actually receives
rtr_t
the task reward: what the team achieved
ctc_t
how much communication was used this step, in bits or messages
λ\lambda
the price. Setting it is a design decision, not a physical constant
Talking is now something the team pays for, so a message has to earn its price.

What λ\lambda does to the learning problem:

  • At λ=0\lambda = 0, a team should talk constantly. What to say becomes uninteresting, since everything can be said.
  • As λ\lambda rises, messages that change no decision get selected against first. This is the mechanism that makes the three tests from Message Content enforceable rather than merely good advice.
  • At high λ\lambda, the team goes quiet and has to coordinate on observations, memory and convention, as in the Coordinate chapter.

Knowledge check

A team is trained with λ = 0 and learns to broadcast its full observation every step. What is the most likely outcome when it is deployed on a channel that charges per bit?

Select one answer.

  • A communication policy πim(⋅∣hti)\pol{\ag}^{m}(\cdot \given h^\ag_t) produces messages from local information, and is learnable like any other policy.
  • Because ∅\varnothing exists, an agent decides both what to say and whether to say anything.
  • A team that can stay silent is more capable, and makes the arrival of a message informative by itself. It also becomes fragile under loss.
  • Recipients may be reached by broadcast, by range, or by targeting.
  • A cost rt′=rt−λctr'_t = r_t - \lambda c_t is what makes efficiency an objective rather than an aspiration. At λ=0\lambda = 0 a team should talk constantly.
  • Constraints must be present during training. A protocol learned for free does not become efficient when charged later.