1.1Coordinate
Let’s say
Section titled “Let’s say”Two agents are working in a kitchen. Their shared goal is to prepare and serve an order as quickly as possible.
- An ingredient is missing.
- Agent A decides to collect it.
- Agent B independently makes the same decision.
- Both agents reach for the same ingredient.
- Meanwhile, nobody prepares the stove.
Both decisions make sense individually. Together, they are inefficient.
shared goal
both reach for one thing
complementary
Same goal. Different joint actions. Different outcomes.
The coordination problem
Section titled “The coordination problem”- The quality of one agent’s action often depends on what the other agents do.
- A shared goal does not automatically produce coordinated behaviour.
- The team receives one reward, so nobody is told what they personally contributed.
In this chapter
Section titled “In this chapter”- Joint decisions: Understand why individual actions must be considered together.
- Independent learning: See what happens when each agent learns its own policy while the others are also learning.
- Centralized and decentralized learning: Compare what information can be used during training and during execution.
- Centralized training with decentralized execution: Learn how agents can use global information while training but still act from local observations at deployment.
- Centralized critics: See how team-level information can improve value estimation during training.
- Credit assignment: Understand why a shared reward does not reveal which individual actions produced the outcome.
- Value decomposition: Learn how VDN and QMIX connect individual action values to a team-level value.
How the ideas build
Section titled “How the ideas build”- Individual decisions
- Joint outcomes
- Team-aware learning
- Decentralized coordination
By the end of this chapter, you should be able to explain why shared rewards are not enough for coordination, and compare different ways of training agents to produce useful joint behaviour.