3.4Training Partner Diversity
Training partner diversity exposes an agent to several cooperative behaviours instead of allowing it to specialize around one fixed partner.
In this section you will
- Contrast fixed-partner training with a partner population
- Describe how population-based training varies interactions
- Argue for behavioural coverage over the raw number of partners
- Frame a partner change as distribution shift
Fixed-Partner Training
Section titled “Fixed-Partner Training”The setup so far.
- the learning agent’s policy
- the single fixed training-partner policy
Agent sees a single pattern of behaviour throughout training. Every regularity in is available to be exploited, and exploiting it is what maximising team return means. The narrow rule from Partner Dependence is not merely permitted here, it is rewarded.
Partner Populations
Section titled “Partner Populations”Instead, draw the partner from a set.
- a population of partners; each episode draws one
Members of can differ in every dimension a convention can form over:
- which roles they take,
- their movement patterns and preferred side of the counter,
- their task preferences, and what they do when idle,
- their communication conventions, if there is a channel,
- their reaction speed, and whether they commit early or wait.
Now the narrow rule stops working. “My partner fetches, so I cook” fails on every population member that also cooks, so the agent has to find something that holds across the variation. Getting the ingredient collected somehow is such a rule; assuming who collects it is not.
Population-Based Training
Section titled “Population-Based Training”The general technique has a name. The textbook describes population-based training as extending training from interaction with a single policy to interaction with a distribution or population of policies, and notes that policy populations can increase the diversity of interactions encountered during training. Murphy’s survey describes the same idea: training against a population of different policies rather than a single one.
There is a substantial literature on how to build and maintain such a population, including self-play variants and methods that grow the population by repeatedly adding best responses. Those are worth knowing exist and are not required here, because the pedagogical content is entirely in what the population is for.
Behavioural Coverage
Section titled “Behavioural Coverage”Here is the mistake worth pre-empting.
| Population | Likely outcome | |
|---|---|---|
| System A | 10 partners from the same run, different seeds | little useful variation |
| System B | 4 partners with genuinely different strategies | real pressure toward generality |
Ten partners that all fetch ingredients present one behavioural pattern ten times. The agent can co-adapt to that pattern exactly as it would to a single partner, and will, because nothing distinguishes the situation from fixed-partner training.
- the population size
- the range of distinct cooperative strategies
This is also why “we trained against 50 partners” is not by itself a generalization claim. The question is what those 50 partners do differently, and answering it requires measuring their behaviour rather than counting them.
Partner Shift as Distribution Shift
Section titled “Partner Shift as Distribution Shift”The population view gives the whole chapter a cleaner statement.
- the policies of everyone other than agent i
- the distribution partners are drawn from while learning
- the distribution actually encountered afterwards
Written this way, partner generalization is distribution shift, and every instinct from that literature transfers. The gap between and is what determines whether an agent will cope, and widening is the direct intervention.
This is how the modern zero-shot coordination literature frames it. ZSC-Eval, which Partner Dependence0 covers, describes zero-shot coordination as an out-of-distribution generalization problem arising from differences between training partners and deployment partners.
Knowledge check
Correct.
Not quite.
Team Y. Its four partners present genuinely different behaviour, so no single convention works across them.
Right. Y’s population applies pressure that X’s does not. X’s 20 seeds may all adopt similar strategies, in which case the agent can co-adapt to that shared pattern exactly as it would to one partner.
Team X. Twenty partners is five times the variation of four.
This counts partners rather than behaviours, which is the mistake the section exists to prevent. Twenty copies of one strategy present one strategy.
Neither, since population size is what matters and both are small.
Size is not what matters. A carefully chosen handful of genuinely distinct partners can apply more pressure toward generality than a large near-identical set.
Team X, because different random seeds guarantee behavioural diversity.
They do not guarantee it. Different seeds sometimes find different conventions and sometimes converge on similar ones, as the Communication Lab showed when two of six runs learned the same protocol. Diversity has to be verified, not assumed.
Explanation
Any population needs its diversity measured rather than asserted, which is exactly the argument of Partner Dependence0.
Training Partner Diversity Summary
Section titled “Training Partner Diversity Summary”- Fixed-partner training rewards co-adaptation: exploiting a partner’s regularities is what maximising team return means.
- Training against a population removes that shortcut, because no single set of habits is safe to assume.
- Population-based training extends learning from one policy to a distribution over policies, increasing the diversity of interactions.
- Number of partners is not behavioural diversity. Ten near-identical partners present one pattern.
- Partner generalization is distribution shift: against .
- Diversity costs peak performance with any single familiar partner. That is the trade being made.