Skip to content
MARL in Cooperative Environments
Edit this page

3.3Partner Generalization

5 min read

Partner generalization is the ability to cooperate with behaviours that were absent from training.

In this section you will

  • Separate the training partner set from the evaluation partner set
  • Define the familiar-to-unseen performance gap
  • Connect that gap to partner dependence
  • Distinguish partner shift from environment shift

Borrow the framing from supervised learning, where it is second nature, and apply it to partners rather than to data.

Figure 1
trained with Πtrainevaluated with ΠtestΠtest⊈Πtrain\text{trained with } \tone{policy}{\Pi_{\text{train}}} \qquad\qquad \text{evaluated with } \tone{conflict}{\Pi_{\text{test}}} \qquad\qquad \tone{conflict}{\Pi_{\text{test}}} \not\subseteq \tone{policy}{\Pi_{\text{train}}}
Πtrain\Pi_{\text{train}}
the set of partners the agent trained with
Πtest\Pi_{\text{test}}
the partners it is evaluated against, containing behaviour it never met
Held-out partners rather than held-out data. Everything else about the setup can stay identical.

The question becomes precise:

How well does this agent cooperate with behaviour that was not present during training?

Note what happens to the usual practice under this framing. Reporting J(πA,πB)J(\pi_A, \pi_B) where πB∈Πtrain\pi_B \in \Pi_{\text{train}} is reporting training performance. Nobody would accept that as a generalization result about a classifier, and it is the standard cooperative MARL number.

A simple diagnostic, useful for building intuition.

Figure 2
Δgen=Jfamiliar⏟with training partners−Junseen⏟with held-out partners\tone{conflict}{\Delta_{\text{gen}}} = \ubrace{policy}{J_{\text{familiar}}}{with training partners} - \ubrace{conflict}{J_{\text{unseen}}}{with held-out partners}
Δgen\Delta_{\text{gen}}
how much performance is lost by changing partner
One number for the distance between looking good and being general.

An example of what it looks like in practice:

PartnerSuccess
Training partner B94%
Training partner C91%
Unseen partner D62%
Unseen partner E48%

Strong in distribution, brittle across partners. Note that averaging all four into “74%” would hide the entire finding, which is the argument for reporting the spread rather than the mean.

The phenomenon needs a name, and it is worth being careful about how firm to make it.

That is a description rather than a quantity. There is no single agreed measurement of “how partner-dependent” a policy is, and treating one as canonical would be misleading. What the field does agree on is the diagnostic procedure: pair agents that did not train together and look at what happens.

Now a distinction that is easy to blur and matters a great deal.

What changesWhat stays the same
Environment generalizationthe kitchen: layout, timings, order mixthe kind of partner
Partner generalizationthe partner’s behaviour and conventionsthe kitchen

An agent can be robust to one and poor at the other, and they call for different fixes.

Environment shift is familiar from single-agent reinforcement learning. The usual remedies are domain randomisation and varied training levels, and the failure mode is a policy that memorised one layout.

Partner shift has no single-agent analogue at all. There is no such thing as a convention held by one agent, so a policy can be perfectly robust to every layout you can generate and still fail the moment its partner is replaced.

Knowledge check

An agent is trained with domain randomisation across hundreds of kitchen layouts, always with the same partner. What does this buy?

Select one answer.

  • Partner generalization is held-out partners, not held-out data: Πtest⊈Πtrain\Pi_{\text{test}} \not\subseteq \Pi_{\text{train}}.
  • Under that framing, the usual reported number is training performance.
  • Δgen=Jfamiliar−Junseen\Delta_{\text{gen}} = J_{\text{familiar}} - J_{\text{unseen}} is a useful teaching diagnostic, not a standard metric. It ignores which partners were used and rewards uniform mediocrity.
  • Report the spread across partners, not the mean. An average hides the finding.
  • Partner dependence is a description of a phenomenon rather than a measured quantity. What is agreed is the procedure: pair agents that did not train together.
  • Environment shift and partner shift are different axes needing different variation. Hold the environment fixed when testing the partner.