3.2Partner Dependence
Adaptation asks whether an agent can continue to cooperate when its partners differ from those encountered during training.
In this section you will
- Describe what changes when a familiar partner is replaced
- Define co-adaptation between agents trained together
- State the limits of familiar-partner performance as evidence
- Identify the conventions an agent has learned about one partner
Partner Change in the Cooperative Kitchen
Section titled “Partner Change in the Cooperative Kitchen”Agent 1 has trained with agent 2 for a very long time. During training, a division of labour settled:
Agent 2 usually collects ingredients, so agent 1 learned to prepare and cook.
Their performance is excellent. Now replace agent 2 with agent 3, trained elsewhere, which learned:
My partner collects ingredients, so I prepare and cook.
Two capable agents, each waiting for the other to fetch the tomato.
The episode begins.
| Step | Agent 1 | Agent 3 |
|---|---|---|
| 1 | waits for the ingredient | waits for the ingredient |
| 2 | waits for the ingredient | waits for the ingredient |
| 3 | waits for the ingredient | waits for the ingredient |
Both agents are individually capable. Both performed excellently during training. Together they cannot serve a single order.
Co-Adaptation
Section titled “Co-Adaptation”Here is the mechanism. When agents train together repeatedly, each policy adapts to the behaviour of the others, and there is nothing to stop that adaptation going further than the task requires.
Agent 1 could have learned the general rule:
When an ingredient is needed, work out who should collect it.
What it actually learned may be much narrower:
Agent 2 collects ingredients, so I prepare.
Both rules produce identical behaviour throughout training. The narrow one is easier to learn, needs no reasoning about who should do what, and is indistinguishable from the general one as long as agent 2 is present.
Limits of Familiar-Partner Performance
Section titled “Limits of Familiar-Partner Performance”Write the thing that was actually optimised.
- the team return of this particular pair of policies
A high tells you that this joint policy is good. It is evidence about one point in the space of possible pairings, and the space is large:
- partners nobody measured
This is the shape of the problem, and it is why the Adapt chapter spends a whole section on evaluation. The standard number reported for a cooperative team is its performance with itself, and that number is systematically optimistic about everything this chapter cares about.
Partner-Specific Conventions
Section titled “Partner-Specific Conventions”Co-adaptation produces conventions: arbitrary
agreements that work because both parties hold them. the Communicate chapter met one kind,
where symbol 2 meant whatever the pair had settled on. Conventions form over
much more than messages:
- roles, who cooks and who fetches,
- task allocation, which agent takes the left side,
- movement, which way round the counter you walk,
- timing, who commits first when both could act,
- message meanings, as in the Communicate chapter.
| Team A + B settled on | Team C + D settled on | |
|---|---|---|
| ingredients | A | D |
| cooking | B | C |
Neither convention is wrong. Nothing about the kitchen makes one correct: the assignment was decided by whichever accident happened first in training. Both teams serve every order.
Pair A with C, and both cook.
Knowledge check
Correct.
Not quite.
That each pair coordinates well together. It establishes nothing about whether any of them can coordinate with an agent it has not trained with.
Correct, and the distinction is the point of the chapter. J(π_A, π_B) is one measurement about one pairing. Partner generalization is a claim about other pairings, and no amount of within-team performance is evidence for it.
That all four have learned general coordination, since they each solved the task.
They each solved the task with a specific partner. The narrow rule "my partner fetches, so I cook" solves the task perfectly and is not general coordination at all.
That the task is easy, since two independent teams both solved it.
The task may well be easy. That is compatible with cross-pairings failing completely, because the difficulty being measured here is compatibility rather than task difficulty.
Nothing, because 95% is not high enough to draw conclusions.
Too dismissive. 95% is strong evidence about the thing it measures, which is this pair working together. The problem is not the size of the number but what the number is about.
Explanation
Asking “with which partner?” of every reported cooperative result is the quickest way to tell a general claim from a specific one.
Partner Dependence Summary
Section titled “Partner Dependence Summary”- Agents trained together co-adapt: each policy adapts to the others, often more narrowly than the task requires.
- A narrow rule such as “my partner fetches, so I cook” is optimal on the training distribution and indistinguishable from general competence while that partner is present.
- measures one pairing. It is the number usually reported, and it is silent about every other partner.
- Co-adaptation produces conventions over roles, allocation, movement, timing and message meanings. Conventions are arbitrary, and separately trained teams settle on different ones.
- Two individually excellent agents can fail together with nothing wrong in either. The Communication Lab measured this: 1.00 within pairs, 0.300 across them.