Skip to content
MARL in Cooperative Environments
Edit this page

3.11Research Connection: ZSC-Eval

5 min read

ZSC-Eval treats the selection of unseen partners as part of the evaluation method rather than an incidental benchmark choice.

In this section you will

  • Distinguish training, evaluation and deployment partner distributions
  • Describe how candidate partners are generated and selected
  • Read the benchmark results and the human study alongside each other
  • State when a held-out partner result supports a generalization claim

Training, Evaluation, and Deployment Distributions

Section titled “Training, Evaluation, and Deployment Distributions”

Training Partner Diversity framed partner generalization as distribution shift between training and deployment. ZSC-Eval sharpens it, and the sharpening is the part worth learning.

The authors note that the significant difference between the deployment-time partners’ distribution and the training partners’ distribution, which is determined by the training algorithm, makes zero-shot coordination a distinctive out-of-distribution generalization challenge. Then they add the observation that does the real work: the potential distribution gap between evaluation partners and deployment-time partners leads to inadequate evaluation, made worse by a lack of appropriate metrics.

Figure 1
ptrainpevalpdeploy\tone{policy}{p_{\text{train}}} \qquad \tone{observe}{p_{\text{eval}}} \qquad \tone{conflict}{p_{\text{deploy}}}
ptrainp_{\text{train}}
partners the agent trained with
pevalp_{\text{eval}}
partners you happened to test with
pdeployp_{\text{deploy}}
partners it will actually meet
Three distributions. Held-out evaluation partners close the first gap and say nothing about the second.

Evaluating Partner Generalization put this as a question about behavioural diversity. ZSC-Eval puts it as a question about approximating the deployment distribution, which is stronger and harder.

Partner Generation, Selection, and Measurement

Section titled “Partner Generation, Selection, and Measurement”

Three components, and the ordering is the argument.

Generate candidate partners using behaviour-preferring rewards, in order to approximate the distribution of deployment-time partners. Note the mechanism: rather than training more partners the ordinary way and hoping they differ, partners are produced deliberately with rewards that prefer particular behaviours.

Select the evaluation partners by Best-Response Diversity (BR-Div). Generating candidates is not enough, because a large set can still be behaviourally narrow. Selection is where diversity is enforced rather than assumed.

Measure generalization with the Best-Response Proximity (BR-Prox) metric, across the selected partners.

The authors benchmark ZSC algorithms in Overcooked and Google Research Football, and report novel empirical findings. They also conduct a human experiment with current ZSC algorithms to verify that ZSC-Eval’s judgements are consistent with human evaluation.

That last step is worth noticing as methodology. A new evaluation instrument needs validating against something outside itself, and pairing agents with people is about the strongest available check that the instrument measures cooperation rather than an artefact of the partner set.

How do you choose evaluation partners that resemble the ones your system will actually meet?

This is a harder question than it sounds, because the honest answer for most projects is that you do not know what pdeployp_{\text{deploy}} looks like. What you can do is stop pretending that partners you generated are a sample from it.

Knowledge check

A team holds out five partners from training, evaluates on them, and reports strong zero-shot coordination. According to this paper's framing, what has still not been established?

Select one answer.

  • Zero-shot coordination is an out-of-distribution problem: the deployment partner distribution differs from the training one, which the training algorithm determines.
  • There are three distributions, not two. Held-out evaluation partners close the train-to-eval gap and leave the eval-to-deploy gap open.
  • ZSC-Eval generates candidate partners with behaviour-preferring rewards, selects them by BR-Div, and measures generalization with BR-Prox.
  • Generation, selection and metric are three separate requirements. Most evaluations supply none.
  • Benchmarked in Overcooked and Google Research Football, with a human experiment to check the instrument against human judgement.

ZSC-Eval: An Evaluation Toolkit and Benchmark for Multi-agent Zero-shot Coordination, Xihuai Wang, Shao Zhang, Wenhao Zhang, Wentao Dong, Jingxiao Chen, Ying Wen and Weinan Zhang. Advances in Neural Information Processing Systems 37 (NeurIPS 2024), Datasets and Benchmarks Track. Proceedings page. The toolkit is released at github.com/sjtu-marl/ZSC-Eval.