Skip to content
MARL in Cooperative Environments
Edit this page

3.8Zero-Shot Coordination

6 min read

Zero-shot coordination evaluates whether an agent can cooperate with an unseen partner without prior joint training, negotiation, or an agreed convention.

In this section you will

  • Define zero-shot coordination
  • Distinguish zero-shot evaluation from a policy that never adapts
  • Explain compatibility in terms of shared conventions
  • List the mechanisms that improve first-encounter coordination

The setup the ZSC literature uses: an ego agent is trained with a limited set of training partners, then evaluated with diverse unseen deployment partners. No joint training, no negotiation, no warm-up round in which the two agree on anything.

“Zero-shot” is about what happened before the encounter. It means no shared history: no co-training, no exchanged conventions, nothing pre-arranged.

Zero-Shot Evaluation and Online Adaptation

Section titled “Zero-Shot Evaluation and Online Adaptation”

Here is the distinction worth getting right.

What it saysWhen it happens
Zero-shot coordinationthe partner was unseen before deploymenta fact about the setup
Online adaptationthe agent uses interaction during deployment to infer how this partner behavesa capability of the agent

These are orthogonal, and they compose. An agent can begin an encounter zero-shot, knowing nothing about its partner, and then adapt as evidence arrives, exactly as in Partner Representations. That is still zero-shot coordination: nothing was arranged in advance.

There is a genuine trade inside this. Adapting takes evidence, and evidence takes steps, and those steps are spent behaving non-committally. On a long task that cost is trivial. On a short one the agent may be better served by a good partner-agnostic prior than by a belief it never has time to sharpen.

Zero-shot coordination fails for a specific reason, and the Communicate chapter already produced it.

Agent A learnedAgent C learned
message 2bring an ingredientbegin cooking

Both learned successful protocols. Paired zero-shot, 2 means one thing to the sender and another to the receiver, and the team does worse than if neither had spoken.

The Communication Lab measured the general case: six independently trained pairs, 1.00 within pairs, 0.300 across them, some pairings at 0.00.

No guarantees here, and it would be dishonest to present these as prescriptions. They are intuitions consistent with the framing, and they are worth holding loosely.

Avoid rigid assumptions. A policy that requires the partner to fetch will fail with any partner that does not. One that checks whether fetching has happened will not.

Respond to observable behaviour. Anything you condition on that the partner actually does is information you genuinely have. Anything you assume is a bet.

Keep more than one interpretation alive. Committing to a partner model after one ambiguous step is how confident-and-wrong happens.

Choose actions that leave room. Where two actions are equally good for you, the one that leaves your partner more options is likely better for the team, because it fails more gracefully if you have misread them.

Prefer interpretable signals. A protocol anchored to something outside the pair has some chance of being read by a stranger, which is exactly the argument of the LangGround connection.

Knowledge check

An agent meets an unfamiliar partner, spends four steps observing it, infers its role preference, and then coordinates well. Is this zero-shot coordination?

Select one answer.

  • Zero-shot coordination: cooperating with an unseen partner at deployment, with no prior coordination. An ego agent trained on limited partners, evaluated on diverse unseen ones.
  • “Zero-shot” constrains what happened before the encounter, not what the agent may infer during it.
  • Online adaptation is a separate axis and composes with it. Beginning zero-shot and then adapting is still zero-shot coordination.
  • Adapting costs steps. On short tasks a good partner-agnostic prior may beat a belief there is no time to sharpen.
  • Compatibility is needed over roles, conventions, communication, action preferences and timing, all of which a co-trained pair gets for free.
  • What tends to help: avoid rigid assumptions, condition on observed behaviour, keep several interpretations alive, leave your partner room, and prefer interpretable signals. These are intuitions, not rules.