2.6Communication Policies
A communication policy determines what an agent sends, whether it sends at a particular time, and which agents receive the message.
In this section you will
- Select what to send from a local observation
- Decide when a message is worth sending
- Choose which agents should receive it
- Price communication inside the shared reward
Message Selection
Section titled “Message Selection”An agent’s message is produced by a policy, exactly like its environment action, and from the same input.
- the communication policy: this agent’s rule for what to say
- its own history, since what is worth saying often depends on what happened earlier
Two things follow from that shape. The communication policy is local: an agent can only report what it knows. And it is learnable: nothing about this is hand-written unless you choose to hand-write it, which is Learning Communication Protocols.
Note the input is rather than . Some of the most useful messages are about change, not state. “The order just changed” needs a memory of the previous order to say at all.
Communication Timing
Section titled “Communication Timing”Because the message space contains , one available message is to send nothing. So the agent faces two questions, and they are genuinely different:
What should I say?
Do I need to say anything at all?
The second is easy to overlook and does most of the work in a constrained system. Most steps in most tasks are unremarkable. If the situation has not changed and the partner’s plan is still right, a message repeating that is pure cost.
That has a sharp edge, and Communication Constraints named it: a protocol where silence means something breaks under message loss. Making silence meaningful buys efficiency and pays for it in fragility.
Message Recipients
Section titled “Message Recipients”Three arrangements, in increasing specificity.
Broadcast. Everyone receives it. Simple, and it scales badly: bandwidth grows with the number of listeners, and a message useful to one partner is noise to the rest.
Range-limited. Only nearby agents receive it, as described in Communication Constraints. Who hears you is determined by the state rather than chosen.
Targeted. The sender picks a recipient. More expressive, and it enlarges the message space, since the agent now chooses both content and address.
Communication Cost
Section titled “Communication Cost”The final piece, and the one that makes the design questions bite.
If sending is free, the answer to “should I say anything?” is always yes: another message can only add information, so a team should talk constantly. That makes the interesting question disappear.
Charging for it brings the question back.
- the reward the team actually receives
- the task reward: what the team achieved
- how much communication was used this step, in bits or messages
- the price. Setting it is a design decision, not a physical constant
What does to the learning problem:
- At , a team should talk constantly. What to say becomes uninteresting, since everything can be said.
- As rises, messages that change no decision get selected against first. This is the mechanism that makes the three tests from Message Content enforceable rather than merely good advice.
- At high , the team goes quiet and has to coordinate on observations, memory and convention, as in the Coordinate chapter.
Knowledge check
Correct.
Not quite.
Performance falls, because the protocol was never selected for efficiency and most of what it sends changes no decision.
Right. With λ = 0 nothing pressured the team to be selective, so the protocol carries whatever was easiest to learn. Charging for that afterwards penalises every irrelevant bit, and no part of training ever identified which bits those were.
Performance is unaffected, since the agents already learned to coordinate well.
They learned to coordinate given unlimited communication. That skill does not transfer to a channel where the same behaviour now subtracts from the reward on every step.
The agents will stop communicating, since messages now cost something.
They would need to re-learn to do that. A deployed policy does not adapt its protocol because the reward function changed; it keeps executing what it learned, now at a cost.
Performance improves, because the cost encourages more efficient messages.
A cost during *training* encourages efficiency. A cost imposed only at deployment just charges for behaviour already fixed. The λ you train with is part of the problem specification.
Explanation
The general lesson: constraints have to be present during training, or the protocol will not be shaped by them.
Communication Policies Summary
Section titled “Communication Policies Summary”- A communication policy produces messages from local information, and is learnable like any other policy.
- Because exists, an agent decides both what to say and whether to say anything.
- A team that can stay silent is more capable, and makes the arrival of a message informative by itself. It also becomes fragile under loss.
- Recipients may be reached by broadcast, by range, or by targeting.
- A cost is what makes efficiency an objective rather than an aspiration. At a team should talk constantly.
- Constraints must be present during training. A protocol learned for free does not become efficient when charged later.