Skip to content
MARL in Cooperative Environments
Edit this page

Part 5: Evaluation

5 min read

An evaluation plan turns the disaster-response design into testable claims. This section asks you to select four to six metrics spanning task performance, coordination, communication, and adaptation, then arrange them across baseline, ablation, stress, and unseen-partner conditions. A final falsification test must state what result would count against the design. The objective is evidence that distinguishes robust cooperation from success under familiar or unusually favourable conditions.

Not twenty. At least one from each category.

What the system exists to do. Survivors reached, missions completed, response time.

Whether the agents actually work together, rather than the task merely getting done. Duplicated search coverage, idle time, conflicting actions, uncovered regions.

What the channel is costing and whether the system depends on it. Messages sent, bandwidth used, performance under message loss.

The most informative of these is the last. A system that degrades gracefully as messages drop is different from one that falls off a cliff, and only a sweep tells you which you have.

Performance with familiar agents, performance with held-out agents, cross-play performance, performance after an agent joins or leaves.

One compact table. Each row is a test condition and what it is there to establish.

Test conditionPurpose
Familiar agentsbaseline
Unseen agentpartner generalization
Message losscommunication robustness
Agent failureadaptation
Higher task loadenvironment robustness

That is the shape. Yours should reflect the design you actually made: if you chose range as your communication constraint rather than loss, test range.

Two disciplines from the Adapt chapter to apply here.

Hold one thing fixed at a time. Vary the partner with the environment fixed, then the environment with the partner fixed. Change both at once and a performance drop tells you something broke, not what.

Say how your held-out agents differ. Held out is a claim about provenance; behaviourally different is a claim about content. A held-out agent generated by re-running your own training with a new seed may behave almost identically, in which case the test measures very little.

The strongest thing you can add, and it is what the Coordinate chapter research connection was for.

What result would convince you that your system learned to coordinate, rather than learned a fixed sequence that happens to work?

The concrete version: train a policy conditioned only on the timestep, ignoring observations entirely, and see how it scores. If it does well, your environment never required coordination and your headline number establishes nothing about it.

Other forms of the same discipline:

  • randomise start positions and task locations so no fixed plan works,
  • hold out scenario layouts entirely,
  • perturb one agent mid-episode and check the others’ behaviour changes,
  • remove an agent and check the team degrades rather than continuing unaffected.

Knowledge check

A submission proposes measuring survivors reached, response time, and messages sent, all with the agents it trained. What is the most important thing missing?

Select one answer.

  • Four to six metrics, at least one per category
  • An evaluation matrix: condition, and what each condition establishes
  • Held-out agents described behaviourally, not just as held out
  • One variable at a time, partner and environment separated
  • One falsification test, and what result would worry you

Next: Deliverables.