3.9Evaluating Partner Generalization
Evaluating partner generalization requires more than reporting performance with a held-out policy.
In this section you will
- Build disjoint training and test partner sets
- Read self-play performance against cross-play performance
- Interpret a cross-play matrix
- Check the behavioural diversity and sample size of the test partners
- Control the environment variation so the comparison is about partners
Disjoint Training and Test Partners
Section titled “Disjoint Training and Test Partners”The first rule, and it is not negotiable.
- partners seen during training
- partners used for evaluation
Disjoint at the level of specific policies, at minimum. Nobody would evaluate a classifier on its training set, and evaluating a cooperative agent with its training partner is the same error with the word “partner” substituted for “example”.
Self-Play and Cross-Play Evidence
Section titled “Self-Play and Cross-Play Evidence”Report for the co-trained pair. It tells you the method can solve the task at all, which is worth knowing, and it is the ceiling against which everything else is read.
It is also the number that says nothing about this chapter, so it cannot be the headline.
Cross-Play Matrices
Section titled “Cross-Play Matrices”The core instrument. Pair every policy with every partner, including ones it never trained with.
Here is a real one. These are the six sender/receiver pairs from the Communicate chapter lab, each trained independently on the same task, scored as percentage success over 4000 evaluation trials.
| R0 | R1 | R2 | R3 | R4 | R5 | ||
|---|---|---|---|---|---|---|---|
| S0 | |||||||
| S1 | |||||||
| S2 | |||||||
| S3 | |||||||
| S4 | |||||||
| S5 |
Read the diagonal first: every entry is 100. That is what each of these agents would report under a standard evaluation, and all six look perfect.
Now the off-diagonal, which averages 30%. Switch the panel to cross-play mean and the six agents separate: S2 and S4 hold up much better than S1 and S3, and nothing in the diagonal predicted that.
Three specific things worth noticing in this matrix.
Some entries are 0%, against a chance baseline of 25%. A confidently mismatched convention is worse than no information, because the receiver acts decisively on a symbol it has misread.
S2 and S4 score 100% with each other. Two independent runs converged on the same convention. Compatibility happens, and neither agent arranged it.
The matrix is symmetric here because the task is, and it usually will not be. Do not assume in general.
Generalization Gap Limitations
Section titled “Generalization Gap Limitations”The gap mode shows familiar-partner score minus cross-play mean, which is the from Partner Generalization.
It is a useful summary and it has a failure mode worth stating plainly: a policy that is uniformly mediocre has a small gap. An agent scoring 40% with everyone has a gap of zero and is not what anybody wanted.
Behavioural Diversity of Test Partners
Section titled “Behavioural Diversity of Test Partners”Now the subtlest requirement, and the one most often missed.
Suppose an agent scores 90% with three unseen partners. Is that strong zero-shot coordination?
Not necessarily. If those three partners happen to behave much like the training partners, then 90% measures very little. The evaluation set was nominally held out and functionally in-distribution.
| Test set | What it actually tests |
|---|---|
| the training partners | nothing about generalization |
| new seeds of the same training algorithm | robustness to seed noise, maybe |
| behaviourally different held-out partners | partner generalization |
The middle row is the trap. Fresh seeds are genuinely held-out policies and can still be near-identical in behaviour, exactly as in the diversity argument of Training Partner Diversity. The Communicate chapter matrix shows both cases at once: most seed pairs disagree, and S2 and S4 do not.
So the question to ask of any evaluation set is:
How different are the behaviours in it?
That requires measuring the partners rather than counting them, which is what Partner Dependence0’s research connection is about.
Partner Sample Size and Summary Statistics
Section titled “Partner Sample Size and Summary Statistics”One pairing is an anecdote. Evaluate over several unseen partners and report
- the mean across partners,
- the variation across them, since a method can be good on average and catastrophic with one common partner type,
- the worst case or lower tail, which is what matters if any single encounter can be costly.
A single averaged number hides precisely the finding this chapter is about. The example table in Partner Generalization averaged to 74% while containing a 48%.
Controlling Environment Variation
Section titled “Controlling Environment Variation”Finally, the discipline from Partner Generalization applied as procedure.
- Fix the environment. Vary only the partner. Any change is partner generalization.
- Then, if you want it, fix the partner and vary the environment. Any change is environment generalization.
- Only then vary both, and read the result knowing what each axis contributes.
Vary both at once from the start and a performance drop tells you something broke, not what.
Knowledge check
Correct.
Not quite.
Evidence that the held-out partners are behaviourally different from the training partners, rather than merely different policy instances.
Exactly. A 3-point gap is either an excellent result or an artefact of an easy test set, and the two are indistinguishable without characterising the partners. Held-out is a claim about provenance; behaviourally different is a claim about content.
A larger gap would be more convincing, since a 3-point gap looks implausible.
A small gap is the desirable outcome, not a red flag by itself. What makes it uninterpretable is not its size but the absence of information about what was tested against.
The absolute numbers are too low to support any conclusion.
88% and 91% are strong scores. The question is what the 88% was measured against, not how large it is.
It should report the environment-generalization gap as well.
Worth having and not the gap here. The claim being made is about partners, so the missing evidence is about partners. Adding a second axis would not repair the first.
Explanation
“Held out from what, and how different?” is the question that separates a generalization result from a well-presented training result.
Partner Generalization Evaluation Summary
Section titled “Partner Generalization Evaluation Summary”- Disjoint train and test partners, or you are reporting training performance.
- Self-play performance is necessary and not sufficient. It is the ceiling, not the headline.
- Cross-play is the instrument: every policy against every partner. A strong agent needs more than a bright diagonal.
- Scores can fall below chance, because a confidently mismatched convention is worse than no information.
- The gap is only meaningful next to a high familiar-partner score, since uniform mediocrity has a small gap.
- The behavioural diversity of the test partners decides what the evaluation measures. Fresh seeds are held out and may not be different.
- Report mean, variation and worst case over several unseen partners.
- Hold the environment fixed first, then vary it separately.