Skip to content
MARL in Cooperative Environments
Edit this page

3.4Training Partner Diversity

5 min read

Training partner diversity exposes an agent to several cooperative behaviours instead of allowing it to specialize around one fixed partner.

In this section you will

  • Contrast fixed-partner training with a partner population
  • Describe how population-based training varies interactions
  • Argue for behavioural coverage over the raw number of partners
  • Frame a partner change as distribution shift

The setup so far.

Figure 1
πA⟷πB\tone{policy}{\pi_A} \longleftrightarrow \tone{policy}{\pi_B}
πA\pi_A
the learning agent’s policy
πB\pi_B
the single fixed training-partner policy
Fixed-partner training repeatedly exposes the learner to one behavioural pattern.

Agent AA sees a single pattern of behaviour throughout training. Every regularity in πB\pi_B is available to be exploited, and exploiting it is what maximising team return means. The narrow rule from Partner Dependence is not merely permitted here, it is rewarded.

Instead, draw the partner from a set.

Figure 2
Π={πB, πC, πD, … }\tone{policy}{\Pi} = \bigl\{\tone{policy}{\pi_B},\ \tone{policy}{\pi_C},\ \tone{policy}{\pi_D},\ \dots\bigr\}
Π\Pi
a population of partners; each episode draws one
Several partners rather than one, so no single set of habits is safe to assume.

Members of Π\Pi can differ in every dimension a convention can form over:

  • which roles they take,
  • their movement patterns and preferred side of the counter,
  • their task preferences, and what they do when idle,
  • their communication conventions, if there is a channel,
  • their reaction speed, and whether they commit early or wait.

Now the narrow rule stops working. “My partner fetches, so I cook” fails on every population member that also cooks, so the agent has to find something that holds across the variation. Getting the ingredient collected somehow is such a rule; assuming who collects it is not.

The general technique has a name. The textbook describes population-based training as extending training from interaction with a single policy to interaction with a distribution or population of policies, and notes that policy populations can increase the diversity of interactions encountered during training. Murphy’s survey describes the same idea: training against a population of different policies rather than a single one.

There is a substantial literature on how to build and maintain such a population, including self-play variants and methods that grow the population by repeatedly adding best responses. Those are worth knowing exist and are not required here, because the pedagogical content is entirely in what the population is for.

Here is the mistake worth pre-empting.

PopulationLikely outcome
System A10 partners from the same run, different seedslittle useful variation
System B4 partners with genuinely different strategiesreal pressure toward generality

Ten partners that all fetch ingredients present one behavioural pattern ten times. The agent can co-adapt to that pattern exactly as it would to a single partner, and will, because nothing distinguishes the situation from fixed-partner training.

Figure 3
number of partners⏟easy to count  ≠  behavioural diversity⏟what actually matters\ubrace{conflict}{\text{number of partners}}{easy to count} \;\neq\; \ubrace{reward}{\text{behavioural diversity}}{what actually matters}
number of partners\text{number of partners}
the population size
behavioural diversity\text{behavioural diversity}
the range of distinct cooperative strategies
Population size alone does not measure behavioural coverage.

This is also why “we trained against 50 partners” is not by itself a generalization claim. The question is what those 50 partners do differently, and answering it requires measuring their behaviour rather than counting them.

The population view gives the whole chapter a cleaner statement.

Figure 3b
training:    π−i∼ptrain(π)deployment:    π−i∼pdeploy(π)\text{training:}\;\; \tone{policy}{\pi_{-\ag}} \sim \tone{policy}{p_{\text{train}}(\pi)} \qquad\qquad \text{deployment:}\;\; \tone{conflict}{\pi_{-\ag}} \sim \tone{conflict}{p_{\text{deploy}}(\pi)}
π−i\pi_{-\ag}
the policies of everyone other than agent i
ptrainp_{\text{train}}
the distribution partners are drawn from while learning
pdeployp_{\text{deploy}}
the distribution actually encountered afterwards
Two distributions over partners. Everything in this chapter is about the distance between them.

Written this way, partner generalization is distribution shift, and every instinct from that literature transfers. The gap between ptrainp_{\text{train}} and pdeployp_{\text{deploy}} is what determines whether an agent will cope, and widening ptrainp_{\text{train}} is the direct intervention.

This is how the modern zero-shot coordination literature frames it. ZSC-Eval, which Partner Dependence0 covers, describes zero-shot coordination as an out-of-distribution generalization problem arising from differences between training partners and deployment partners.

Knowledge check

Two teams train an agent against a population. Team X uses 20 partners generated by re-running the same algorithm with different seeds. Team Y uses 4 partners: one ingredient-first, one cooking-first, one that fills whichever role is empty, and one that waits before committing. Which is more likely to generalise?

Select one answer.

  • Fixed-partner training rewards co-adaptation: exploiting a partner’s regularities is what maximising team return means.
  • Training against a population Π\Pi removes that shortcut, because no single set of habits is safe to assume.
  • Population-based training extends learning from one policy to a distribution over policies, increasing the diversity of interactions.
  • Number of partners is not behavioural diversity. Ten near-identical partners present one pattern.
  • Partner generalization is distribution shift: ptrainp_{\text{train}} against pdeployp_{\text{deploy}}.
  • Diversity costs peak performance with any single familiar partner. That is the trade being made.