Skip to content
MARL in Cooperative Environments
Edit this page

2.11Communication Lab

7 min read

The Communication Lab is a hands-on Colab investigation of the smallest task in which one agent must communicate private information to another. Six experiments compare performance without a channel, learn a four-symbol protocol, restrict its capacity, introduce message loss, inspect the learned symbol mapping, and test cross-play with an independently trained partner. The provided tabular learners run on free Colab CPU, keeping the focus on measured communication behaviour and interpretation rather than implementation scale.

01_communicate.ipynb
Time
about 30 minutes
Compute
CPU is enough; no GPU, nothing to install
  • An order arrives, one of four dishes. Only agent 1 can see which.
  • Only agent 2 can act. It picks one of four preparations.
  • The team scores +1+1 if the preparation matches the dish, 00 otherwise.
  • Between them, a channel carrying one symbol from a message space M\mathcal{M}.

Agent 1 knows and cannot act. Agent 2 acts and cannot know. The asymmetry from Communication in Cooperative MARL, reduced until nothing else is left.

Both agents are plain tabular learners. The sender learns a value for (order,message)(\text{order}, \text{message}); the receiver learns a value for (message received,action)(\text{message received}, \text{action}). Neither is told what any symbol means, because there is nothing to tell: the symbols have no meaning until the pair gives them one.

A message space of size 1 is a channel that can only say one thing, which is the same as no channel: the symbol never varies, so it carries nothing.

Measured success: 0.25. Random guessing over four dishes is 1/41/4.

Note what agent 2 is not: badly trained. No policy it could learn would do better, because nothing reaching it depends on the order. This is the baseline every later number should be read against.

Widen the channel to four symbols, enough to name every dish, and change nothing else.

Measured success: 1.00.

The agents worked out an assignment between dishes and symbols with no dictionary, no supervision, and no shared initialisation. That is the whole of what “learning a protocol” means.

Vary ∣M∣|\mathcal{M}| and hold everything else fixed. Ten seeds each.

Task success against Size of the message space. 1: 0.3, 2: 0.5, 3: 0.8, 4: 1, 8: 1 00.30.50.8112348Size of the message spaceTask success

∥M∥\|\mathcal{M}\|12348
Success0.2500.5000.7501.0001.000

Read the numbers rather than the curve, because they are exact. With ∣M∣|\mathcal{M}| symbols the team names ∣M∣|\mathcal{M}| of the four dishes precisely and guesses the rest, giving ∣M∣/4|\mathcal{M}|/4 until the channel is wide enough, then nothing further.

This is Communication Constraints made measurable. A narrow channel does not make the team describe each situation more briefly. It forces the team to merge situations, and every merged pair costs exactly a coin flip.

Four symbols, and now the channel drops a fraction of them. A dropped message arrives as the empty message, indistinguishable from silence.

plossp_{\text{loss}}MeasuredPredicted by 1−34p1 - \tfrac{3}{4}p
0.01.0001.000
0.10.9240.925
0.30.7780.775
0.50.6090.625

The prediction is worth deriving. When the message arrives the team is right; when it is lost, agent 2 is back to guessing at 1/41/4. So expected success is (1−p)⋅1+p⋅14=1−34p(1-p) \cdot 1 + p \cdot \tfrac{1}{4} = 1 - \tfrac{3}{4}p.

That measurement and prediction agree tells you something specific: the agents learned nothing clever about loss. They learned the same protocol as before and simply lose it a fraction of the time. A team that had adapted to an unreliable channel would beat this line.

The team scores 1.00, so it has agreed on something. Here is what, from seed 0:

DishSymbol sentReceiver prepares
010
121
232
303

Look at the dish-to-symbol column. It is not the identity, and nothing is wrong with that.

This is Learning Communication Protocols on one screen. The symbol 1 does not mean dish 0 in any intrinsic sense. It means dish 0 because this receiver prepares dish 0 on seeing it. Any permutation works equally well, provided both agents permute together.

The notebook asks you to re-run with other seeds and count how many distinct protocols you find.

Six independent training runs. Each pair scores 1.00 with itself. Now take the sender from one run and the receiver from another.

R0R1R2R3R4R5
S01.000.250.500.000.500.25
S10.251.000.000.500.000.00
S20.500.001.000.251.000.50
S30.000.500.251.000.250.00
S40.500.001.000.251.000.50
S50.250.000.500.000.501.00

Diagonal mean: 1.000. Off-diagonal mean: 0.300.

Three things to notice.

The diagonal is what these agents would report if evaluated with the partner they trained with. Every one of them looks perfect.

Some off-diagonal entries are 0.00, which is worse than guessing. A confidently wrong protocol is more damaging than no protocol, because the receiver acts decisively on a symbol it has misunderstood.

S2 and S4 score 1.00 with each other. Those two runs happened to converge on the same assignment. Compatibility is possible; it is just not something either agent arranged.

Knowledge check

In the matrix above, some off-diagonal cells are 0.00 while random guessing would give 0.25. How can a trained pair do worse than chance?

Select one answer.

  • With no usable channel, agent 2 guesses at 1/41/4, and no amount of training helps.
  • With a wide enough channel, two tabular learners with no dictionary reach 1.00.
  • Capacity merges situations. Success is ∣M∣/4|\mathcal{M}|/4 until the channel is wide enough, and nothing more after that.
  • Loss degrades as 1−34p1 - \tfrac{3}{4}p here, which is the signature of a protocol that ignores the channel rather than adapting to it.
  • Symbol assignments are arbitrary. Seed 0’s mapping is a permutation, and it works.
  • Cross-play collapses from 1.000 to 0.300, with some pairings worse than chance. Nothing is wrong with any agent; the agreement is missing.

Run it yourself:Open in Colab

Two individually capable agents suddenly stop cooperating. The task did not change, the channel did not change, and neither agent got worse.

So what, exactly, did each of them learn?

That is the Adapt chapter.