3.3Partner Generalization
Partner generalization is the ability to cooperate with behaviours that were absent from training.
In this section you will
- Separate the training partner set from the evaluation partner set
- Define the familiar-to-unseen performance gap
- Connect that gap to partner dependence
- Distinguish partner shift from environment shift
Training and Evaluation Partners
Section titled “Training and Evaluation Partners”Borrow the framing from supervised learning, where it is second nature, and apply it to partners rather than to data.
- the set of partners the agent trained with
- the partners it is evaluated against, containing behaviour it never met
The question becomes precise:
How well does this agent cooperate with behaviour that was not present during training?
Note what happens to the usual practice under this framing. Reporting where is reporting training performance. Nobody would accept that as a generalization result about a classifier, and it is the standard cooperative MARL number.
The Partner Generalization Gap
Section titled “The Partner Generalization Gap”A simple diagnostic, useful for building intuition.
- how much performance is lost by changing partner
An example of what it looks like in practice:
| Partner | Success |
|---|---|
| Training partner B | 94% |
| Training partner C | 91% |
| Unseen partner D | 62% |
| Unseen partner E | 48% |
Strong in distribution, brittle across partners. Note that averaging all four into “74%” would hide the entire finding, which is the argument for reporting the spread rather than the mean.
Generalization Gap and Partner Dependence
Section titled “Generalization Gap and Partner Dependence”The phenomenon needs a name, and it is worth being careful about how firm to make it.
That is a description rather than a quantity. There is no single agreed measurement of “how partner-dependent” a policy is, and treating one as canonical would be misleading. What the field does agree on is the diagnostic procedure: pair agents that did not train together and look at what happens.
Partner Shift and Environment Shift
Section titled “Partner Shift and Environment Shift”Now a distinction that is easy to blur and matters a great deal.
| What changes | What stays the same | |
|---|---|---|
| Environment generalization | the kitchen: layout, timings, order mix | the kind of partner |
| Partner generalization | the partner’s behaviour and conventions | the kitchen |
An agent can be robust to one and poor at the other, and they call for different fixes.
Environment shift is familiar from single-agent reinforcement learning. The usual remedies are domain randomisation and varied training levels, and the failure mode is a policy that memorised one layout.
Partner shift has no single-agent analogue at all. There is no such thing as a convention held by one agent, so a policy can be perfectly robust to every layout you can generate and still fail the moment its partner is replaced.
Knowledge check
Correct.
Not quite.
Environment generalization, and nothing about partner generalization. The partner never varied, so no pressure was ever applied to it.
Exactly. Randomising the environment produces robustness to environments. Partner conventions were held fixed across all of that training, so the agent had every opportunity to co-adapt to one partner and no reason not to.
Both, since a policy robust to environment variation is robust in general.
Robustness does not generalise across kinds of shift. The two require different variation during training, and there is no reason varying one would cover the other.
Partner generalization, because varied layouts force the agents to renegotiate roles.
Appealing, and usually false. The pair can carry the same convention across every layout, and if the convention works everywhere then varied layouts reinforce it rather than loosening it.
Neither, since domain randomisation only helps in simulation.
Too dismissive. Domain randomisation genuinely produces environment robustness that transfers. The limitation here is narrower: it is the wrong axis of variation for this problem.
Explanation
“Which axis did you actually vary?” is the question that predicts which generalization you get.
Partner Generalization Summary
Section titled “Partner Generalization Summary”- Partner generalization is held-out partners, not held-out data: .
- Under that framing, the usual reported number is training performance.
- is a useful teaching diagnostic, not a standard metric. It ignores which partners were used and rewards uniform mediocrity.
- Report the spread across partners, not the mean. An average hides the finding.
- Partner dependence is a description of a phenomenon rather than a measured quantity. What is agreed is the procedure: pair agents that did not train together.
- Environment shift and partner shift are different axes needing different variation. Hold the environment fixed when testing the partner.