Skip to content
MARL in Cooperative Environments
Edit this page

3.1Adapt

2 min read

Agent A and Agent B have trained together for thousands of episodes. They have developed an effective routine:

  • Agent B usually collects ingredients.
  • Agent A usually prepares and cooks.
  • They rarely duplicate work.
  • Their team reward is consistently high.

Now replace Agent B with Agent C. Agent C is also capable, but it learned a different routine:

  • Agent C expects its partner to collect ingredients.
  • Agent C prepares to cook.

The next order arrives.

  • Agent A waits for Agent C to collect the ingredient.
  • Agent C waits for Agent A.
  • Neither moves.

Both agents were successful with their original partners. Together, they fail.

order served

familiar partner

nobody fetches

new partner

Same task. Same shared reward. Different partner.

  • Strong performance with a familiar partner does not mean an agent learned a general ability to cooperate.
  • Agents become dependent on habits, roles, conventions and communication protocols developed during training.
  • Nothing is broken in either agent. What is missing is the agreement between them.
  • Partner dependence: Understand how agents become specialized to familiar partners.
  • Partner generalization: Compare performance with familiar and unfamiliar partners.
  • Training partner diversity: Learn why exposure to different behaviours can improve robustness.
  • Agent modelling: See how an agent can predict what another agent is likely to do.
  • Partner representations: Explore how useful behavioural information can be summarized compactly.
  • Ad hoc teamwork: Learn how agents cooperate with partners they did not design or train with.
  • Zero-shot coordination: Study cooperation with unseen partners without prior joint training.
  • Evaluation: Use cross-play and held-out partners to determine whether cooperation really generalizes.
  • N-Agent Ad Hoc Teamwork and ZSC-Eval: Connect the chapter to recent research on partner generalization and evaluation.
  1. Familiar partners
  2. Partner diversity
  3. Agent modelling
  4. Unseen partners
  5. Evaluation

And the loop an adaptive agent runs while it plays

  1. Observe
  2. Infer
  3. Adjust
  4. Cooperate

By the end of this chapter, you should be able to distinguish specialization from general cooperation, and evaluate whether an agent can work effectively with unfamiliar partners.