2.8Interpretable Communication
Interpretable communication uses messages whose meaning can be understood outside the particular agents that learned them.
In this section you will
- Recognize a protocol that works and cannot be read
- State what interpretability buys for inspection, safety and transfer
- Separate a message symbol from the convention that gives it meaning
- Describe what it means to ground a protocol externally
Effective but Opaque Protocols
Section titled “Effective but Opaque Protocols”Suppose agents 1 and 2 have converged. Whenever an ingredient is needed,
agent 1 sends 2, and agent 2 fetches one. The team serves every order. By
any task metric, the communication is a success.
Now watch them work. What you see is:
2
That is all. The symbol is an integer with no intrinsic content, and the semantics live in the pair of policies rather than anywhere you can read.
- improves the shared task return
- has semantics recoverable outside the co-trained pair
With continuous messages it is worse. A 16-dimensional real vector that reliably conveys the right thing is not inspectable at all, and its structure need not correspond to anything a person would name.
Benefits of Interpretability
Section titled “Benefits of Interpretability”Three reasons, and only the third is about coordination.
Debugging. When a team fails, an unreadable protocol gives you nothing to inspect. You can see that coordination broke and not why.
Trust and oversight. A deployed system whose agents exchange unreadable signals cannot be audited. In a safety-relevant setting this alone can rule the design out.
Interoperability. This is the one that turns out to matter most, and it takes the rest of the section.
Symbols and Protocol Conventions
Section titled “Symbols and Protocol Conventions”Learning Communication Protocols established that symbol labels are arbitrary: any permutation works, provided both agents permute together.
Now take that seriously. Two teams, trained separately, on the same task.
| Team A+B learned | Team C+D learned | |
|---|---|---|
2 | bring an ingredient | serve the dish |
3 | serve the dish | bring an ingredient |
Both protocols work. Both teams serve every order. Neither is wrong, and there is no fact about the task that makes one correct: the assignment was settled by whichever accident happened first during training.
Now pair agent A with agent D.
A sends 2, meaning bring an ingredient. D receives 2, and serves an
empty plate.
This is worth comparing with the equivalent problem in the Coordinate chapter. There, the difficulty was that agents choosing separately might not choose compatibly on a given step. Here the difficulty is that agents might not even share the vocabulary in which compatibility could be expressed.
External Semantic Grounding
Section titled “External Semantic Grounding”A grounded protocol is one whose symbols are tied to something outside the pair that invented them: a natural-language meaning, a task-level referent, a shared prior. Grounding costs something, since it constrains what the agents may learn. What it buys is a protocol that another agent has some chance of reading.
That trade is exactly what the next section’s research connection studies, and the failure above is exactly the problem the Adapt chapter opens with. The bridge between the two chapters is this sentence: a protocol learned with one partner may not survive a different one.
Knowledge check
Correct.
Not quite.
Each team learned a working convention, and the conventions differ. Neither agent is faulty; the shared agreement is what is missing.
Exactly. Symbol assignments are arbitrary, so separate training runs settle on different ones. Both protocols are internally valid and mutually unintelligible, which is why cross-team performance collapses while within-team performance is high.
One of the two protocols must be worse, and the cross-play result reveals which.
Both scored 95% with their own partners, so neither is worse at the task. Cross-play measures compatibility, not quality, and the two can be completely independent.
The agents overfitted to the environment and need more training data.
The environment is the same in both cases; only the partner changed. More training with the same partner would strengthen the convention rather than loosen it, making the problem worse.
The message space is too small to express what both teams need.
Capacity is not the issue: each team expressed everything it needed within the same space. The issue is that they assigned the available symbols differently.
Explanation
“Both work, and they disagree” is a failure mode with no single-agent analogue, and it is why the Adapt chapter exists.
Interpretable Communication Summary
Section titled “Interpretable Communication Summary”- Effective communication does not imply interpretable communication. A protocol can be perfect on the task and unreadable to anyone.
- Continuous messages are worse in this respect than small discrete ones, which at least can be tabulated.
- Noticing that a symbol correlates with a situation is not sufficient evidence that the symbol means that situation.
- Interpretability buys debugging, oversight, and above all interoperability.
- Symbol assignments are arbitrary, so separately trained teams learn incompatible conventions. Two individually excellent agents can fail together with nothing wrong in either.
- A protocol is a convention held by a pair, not a property of the task. That is the bridge to the Adapt chapter.