2.11Communication Lab
The Communication Lab is a hands-on Colab investigation of the smallest task in which one agent must communicate private information to another. Six experiments compare performance without a channel, learn a four-symbol protocol, restrict its capacity, introduce message loss, inspect the learned symbol mapping, and test cross-play with an independently trained partner. The provided tabular learners run on free Colab CPU, keeping the focus on measured communication behaviour and interpretation rather than implementation scale.
Experimental Task
Section titled “Experimental Task”- An order arrives, one of four dishes. Only agent 1 can see which.
- Only agent 2 can act. It picks one of four preparations.
- The team scores if the preparation matches the dish, otherwise.
- Between them, a channel carrying one symbol from a message space .
Agent 1 knows and cannot act. Agent 2 acts and cannot know. The asymmetry from Communication in Cooperative MARL, reduced until nothing else is left.
Both agents are plain tabular learners. The sender learns a value for ; the receiver learns a value for . Neither is told what any symbol means, because there is nothing to tell: the symbols have no meaning until the pair gives them one.
Experiment 1 · No channel
Section titled “Experiment 1 · No channel”A message space of size 1 is a channel that can only say one thing, which is the same as no channel: the symbol never varies, so it carries nothing.
Measured success: 0.25. Random guessing over four dishes is .
Note what agent 2 is not: badly trained. No policy it could learn would do better, because nothing reaching it depends on the order. This is the baseline every later number should be read against.
Experiment 2 · Four symbols
Section titled “Experiment 2 · Four symbols”Widen the channel to four symbols, enough to name every dish, and change nothing else.
Measured success: 1.00.
The agents worked out an assignment between dishes and symbols with no dictionary, no supervision, and no shared initialisation. That is the whole of what “learning a protocol” means.
Experiment 3 · Capacity
Section titled “Experiment 3 · Capacity”Vary and hold everything else fixed. Ten seeds each.
| 1 | 2 | 3 | 4 | 8 | |
|---|---|---|---|---|---|
| Success | 0.250 | 0.500 | 0.750 | 1.000 | 1.000 |
Read the numbers rather than the curve, because they are exact. With symbols the team names of the four dishes precisely and guesses the rest, giving until the channel is wide enough, then nothing further.
This is Communication Constraints made measurable. A narrow channel does not make the team describe each situation more briefly. It forces the team to merge situations, and every merged pair costs exactly a coin flip.
Experiment 4 · Message loss
Section titled “Experiment 4 · Message loss”Four symbols, and now the channel drops a fraction of them. A dropped message arrives as the empty message, indistinguishable from silence.
| Measured | Predicted by | |
|---|---|---|
| 0.0 | 1.000 | 1.000 |
| 0.1 | 0.924 | 0.925 |
| 0.3 | 0.778 | 0.775 |
| 0.5 | 0.609 | 0.625 |
The prediction is worth deriving. When the message arrives the team is right; when it is lost, agent 2 is back to guessing at . So expected success is .
That measurement and prediction agree tells you something specific: the agents learned nothing clever about loss. They learned the same protocol as before and simply lose it a fraction of the time. A team that had adapted to an unreliable channel would beat this line.
Experiment 5 · Read the protocol
Section titled “Experiment 5 · Read the protocol”The team scores 1.00, so it has agreed on something. Here is what, from seed 0:
| Dish | Symbol sent | Receiver prepares |
|---|---|---|
| 0 | 1 | 0 |
| 1 | 2 | 1 |
| 2 | 3 | 2 |
| 3 | 0 | 3 |
Look at the dish-to-symbol column. It is not the identity, and nothing is wrong with that.
This is Learning Communication Protocols on one screen. The symbol 1 does
not mean dish 0 in any intrinsic sense. It means dish 0 because this receiver
prepares dish 0 on seeing it. Any permutation works equally well, provided
both agents permute together.
The notebook asks you to re-run with other seeds and count how many distinct protocols you find.
Experiment 6 · Pair them with a stranger
Section titled “Experiment 6 · Pair them with a stranger”Six independent training runs. Each pair scores 1.00 with itself. Now take the sender from one run and the receiver from another.
| R0 | R1 | R2 | R3 | R4 | R5 | |
|---|---|---|---|---|---|---|
| S0 | 1.00 | 0.25 | 0.50 | 0.00 | 0.50 | 0.25 |
| S1 | 0.25 | 1.00 | 0.00 | 0.50 | 0.00 | 0.00 |
| S2 | 0.50 | 0.00 | 1.00 | 0.25 | 1.00 | 0.50 |
| S3 | 0.00 | 0.50 | 0.25 | 1.00 | 0.25 | 0.00 |
| S4 | 0.50 | 0.00 | 1.00 | 0.25 | 1.00 | 0.50 |
| S5 | 0.25 | 0.00 | 0.50 | 0.00 | 0.50 | 1.00 |
Diagonal mean: 1.000. Off-diagonal mean: 0.300.
Three things to notice.
The diagonal is what these agents would report if evaluated with the partner they trained with. Every one of them looks perfect.
Some off-diagonal entries are 0.00, which is worse than guessing. A confidently wrong protocol is more damaging than no protocol, because the receiver acts decisively on a symbol it has misunderstood.
S2 and S4 score 1.00 with each other. Those two runs happened to converge on the same assignment. Compatibility is possible; it is just not something either agent arranged.
Knowledge check
Correct.
Not quite.
The receiver acts decisively on a symbol it has misunderstood, so a systematic mismatch maps every dish to a consistently wrong preparation.
Exactly. Guessing gets a quarter right by luck. A mismatched protocol is not guessing: it is a reliable mapping composed with the wrong dictionary, so it can be reliably wrong on every dish. Confident misunderstanding is worse than no information.
The agents were undertrained, so the 0.00 entries reflect incomplete learning.
Every one of these agents scores 1.00 with its own partner, which is only possible if training converged. The zeros are a property of the pairing, not of either policy.
The evaluation has too few trials, so 0.00 is noise.
The evaluation runs 4000 trials. A 0.00 over that many trials is a systematic mapping to wrong answers, not sampling error.
Message loss corrupted the cross-pairings.
This experiment runs with no loss. The only thing that changed between the diagonal and the off-diagonal is which receiver the sender was paired with.
Explanation
“Worse than chance” is the signature of a confidently mismatched convention, and it is a useful thing to be able to recognise.
Experimental Findings
Section titled “Experimental Findings”- With no usable channel, agent 2 guesses at , and no amount of training helps.
- With a wide enough channel, two tabular learners with no dictionary reach 1.00.
- Capacity merges situations. Success is until the channel is wide enough, and nothing more after that.
- Loss degrades as here, which is the signature of a protocol that ignores the channel rather than adapting to it.
- Symbol assignments are arbitrary. Seed 0’s mapping is a permutation, and it works.
- Cross-play collapses from 1.000 to 0.300, with some pairings worse than chance. Nothing is wrong with any agent; the agreement is missing.
Cross-Play Implications
Section titled “Cross-Play Implications”Two individually capable agents suddenly stop cooperating. The task did not change, the channel did not change, and neither agent got worse.
So what, exactly, did each of them learn?
That is the Adapt chapter.