Skip to content
MARL in Cooperative Environments
Edit this page

1.11Research Connection: SMACv2

5 min read

SMACv2 illustrates why evaluating coordination requires environments that make policies respond to observations rather than replay fixed routines.

In this section you will

  • Describe what the original SMAC benchmark evaluates
  • Explain the open-loop policy problem the SMACv2 authors identified
  • Examine how randomized teams and starting conditions change the test
  • Apply the diagnostic to a benchmark of your own

SMAC, the StarCraft Multi-Agent Challenge, became the standard testbed for cooperative multi-agent reinforcement learning. Each unit in a battle is an agent with its own local view, choosing its own actions, with the team sharing one reward. Structurally it is a Dec-POMDP, and it is the environment most CTDE methods, QMIX included, were evaluated on.

That much of the StarCraft is all you need. The mechanics are not the point.

Open-Loop Policies as a Benchmark Diagnostic

Section titled “Open-Loop Policies as a Benchmark Diagnostic”

The concern is about what a benchmark requires, and it has a sharp test.

Suppose an agent ignores its observations entirely and conditions only on the timestep, step 1, do this; step 2, do that. Call that an open-loop policy: it does not react to anything. If such a policy scores well, then the task never actually required reacting, and success on it is not evidence of coordination.

The SMACv2 authors ran that test. On SMAC scenarios, policies conditioned only on the timestep turned out to be competitive, which indicates the benchmark carried too little variation from episode to episode. Teams could learn a sequence rather than a strategy and still win.

The fix is to make episodes genuinely vary, so that a fixed sequence cannot work. SMACv2 randomises team compositions and starting positions across episodes, and revises the unit sight and attack range mechanics.

The consequence is what matters here: a policy now has to condition on what it observes, because the situation is not the same one it saw last time. That makes the benchmark a test of closed-loop coordination, coordination that reacts, rather than of sequence memorisation.

Everything you have learned is what the benchmark is testing:

  • joint actions, outcomes depend on the whole team’s choices,
  • partial observability, each unit sees locally,
  • decentralized policies, each agent acts on its own observation,
  • CTDE, training may use the global state; execution may not,
  • value decomposition, how QMIX and VDN produce local action selection,
  • coordination under changing states, the thing SMACv2 was built to demand.

And the frontier question it leaves you with:

Can a benchmark distinguish genuine adaptive coordination from policies that exploit predictable environment structure?

Hold onto that. the Adapt chapter asks the same question with the partner varying instead of the environment, and it turns out to be the harder version.

Knowledge check

Why is it damaging for a coordination benchmark if an open-loop policy, one conditioned only on the timestep, scores well?

Select one answer.

  • A high benchmark score is evidence about the benchmark as much as about the method.
  • An open-loop policy conditions only on the timestep. If it scores well, the task did not require reacting to anything.
  • The SMACv2 authors found such policies competitive on SMAC, indicating too little episode-to-episode variation.
  • SMACv2 randomises team compositions and starting positions and revises sight and attack ranges, so policies must condition on what they observe.
  • Better methods and better measurement are different kinds of progress.

SMACv2: An Improved Benchmark for Cooperative Multi-Agent Reinforcement Learning, Benjamin Ellis, Jonathan Cook, Skander Moalla, Mikayel Samvelyan, Mingfei Sun, Anuj Mahajan, Jakob N. Foerster and Shimon Whiteson (NeurIPS 2023 Datasets and Benchmarks Track, arXiv:2212.07489).

The original benchmark is The StarCraft Multi-Agent Challenge, Mikayel Samvelyan et al. (2019, arXiv:1902.04043).