Skip to content
MARL in Cooperative Environments
Edit this page

1.1Coordinate

2 min read

Two agents are working in a kitchen. Their shared goal is to prepare and serve an order as quickly as possible.

  • An ingredient is missing.
  • Agent A decides to collect it.
  • Agent B independently makes the same decision.
  • Both agents reach for the same ingredient.
  • Meanwhile, nobody prepares the stove.

Both decisions make sense individually. Together, they are inefficient.

shared goal

both reach for one thing

complementary

Same goal. Different joint actions. Different outcomes.

  • The quality of one agent’s action often depends on what the other agents do.
  • A shared goal does not automatically produce coordinated behaviour.
  • The team receives one reward, so nobody is told what they personally contributed.
  • Joint decisions: Understand why individual actions must be considered together.
  • Independent learning: See what happens when each agent learns its own policy while the others are also learning.
  • Centralized and decentralized learning: Compare what information can be used during training and during execution.
  • Centralized training with decentralized execution: Learn how agents can use global information while training but still act from local observations at deployment.
  • Centralized critics: See how team-level information can improve value estimation during training.
  • Credit assignment: Understand why a shared reward does not reveal which individual actions produced the outcome.
  • Value decomposition: Learn how VDN and QMIX connect individual action values to a team-level value.
  1. Individual decisions
  2. Joint outcomes
  3. Team-aware learning
  4. Decentralized coordination

By the end of this chapter, you should be able to explain why shared rewards are not enough for coordination, and compare different ways of training agents to produce useful joint behaviour.