Skip to content
MARL in Cooperative Environments
Edit this page

3.2Partner Dependence

5 min read

Adaptation asks whether an agent can continue to cooperate when its partners differ from those encountered during training.

In this section you will

  • Describe what changes when a familiar partner is replaced
  • Define co-adaptation between agents trained together
  • State the limits of familiar-partner performance as evidence
  • Identify the conventions an agent has learned about one partner

Agent 1 has trained with agent 2 for a very long time. During training, a division of labour settled:

Agent 2 usually collects ingredients, so agent 1 learned to prepare and cook.

Their performance is excellent. Now replace agent 2 with agent 3, trained elsewhere, which learned:

My partner collects ingredients, so I prepare and cook.

Two robot agents working in one kitchen. The left agent is at a chopping board, the right agent at a stove with a pot. They share one counter. Agent 1Agent 2

Two capable agents, each waiting for the other to fetch the tomato.

The episode begins.

StepAgent 1Agent 3
1waits for the ingredientwaits for the ingredient
2waits for the ingredientwaits for the ingredient
3waits for the ingredientwaits for the ingredient

Both agents are individually capable. Both performed excellently during training. Together they cannot serve a single order.

Here is the mechanism. When agents train together repeatedly, each policy adapts to the behaviour of the others, and there is nothing to stop that adaptation going further than the task requires.

Agent 1 could have learned the general rule:

When an ingredient is needed, work out who should collect it.

What it actually learned may be much narrower:

Agent 2 collects ingredients, so I prepare.

Both rules produce identical behaviour throughout training. The narrow one is easier to learn, needs no reasoning about who should do what, and is indistinguishable from the general one as long as agent 2 is present.

Write the thing that was actually optimised.

Figure 1
J(πA, πB)\tone{policy}{J\bigl(\pi_A,\ \pi_B\bigr)}
J(πA,πB)J(\pi_A, \pi_B)
the team return of this particular pair of policies
One number about one pairing. It says nothing about any other pairing.

A high J(πA,πB)J(\pi_A, \pi_B) tells you that this joint policy is good. It is evidence about one point in the space of possible pairings, and the space is large:

Figure 2
J(πA,πC),J(πA,πD),J(πA,πE),…J(\pi_A, \tone{conflict}{\pi_C}), \quad J(\pi_A, \tone{conflict}{\pi_D}), \quad J(\pi_A, \tone{conflict}{\pi_E}), \quad \dots
πC,πD,πE\pi_C, \pi_D, \pi_E
partners nobody measured
Every other pairing agent A might ever be put in. None of them was measured.

This is the shape of the problem, and it is why the Adapt chapter spends a whole section on evaluation. The standard number reported for a cooperative team is its performance with itself, and that number is systematically optimistic about everything this chapter cares about.

Co-adaptation produces conventions: arbitrary agreements that work because both parties hold them. the Communicate chapter met one kind, where symbol 2 meant whatever the pair had settled on. Conventions form over much more than messages:

  • roles, who cooks and who fetches,
  • task allocation, which agent takes the left side,
  • movement, which way round the counter you walk,
  • timing, who commits first when both could act,
  • message meanings, as in the Communicate chapter.
Team A + B settled onTeam C + D settled on
ingredientsAD
cookingBC

Neither convention is wrong. Nothing about the kitchen makes one correct: the assignment was decided by whichever accident happened first in training. Both teams serve every order.

Pair A with C, and both cook.

Knowledge check

Two teams each reach 95% success with their training partners. What does this establish about the four agents' ability to coordinate?

Select one answer.

  • Agents trained together co-adapt: each policy adapts to the others, often more narrowly than the task requires.
  • A narrow rule such as “my partner fetches, so I cook” is optimal on the training distribution and indistinguishable from general competence while that partner is present.
  • J(πA,πB)J(\pi_A, \pi_B) measures one pairing. It is the number usually reported, and it is silent about every other partner.
  • Co-adaptation produces conventions over roles, allocation, movement, timing and message meanings. Conventions are arbitrary, and separately trained teams settle on different ones.
  • Two individually excellent agents can fail together with nothing wrong in either. The Communication Lab measured this: 1.00 within pairs, 0.300 across them.