1.13Coordinate Worksheet
This five-part digital activity consolidates the chapter’s central coordination ideas. You will first match learning approaches to their applications, then calculate a joint-action-space size and a VDN team value. An interactive graph shows how joint-action counts grow with the number of agents. The final prompt asks you to weigh the training benefits of CTDE against its additional complexity. Every item except that trade-off response is automatically checkable.
What you will be able to do
- Match common MARL training approaches to the problems they address. Understand
- Calculate the size and value of simple joint-action problems. Apply
- Interpret how coordination difficulty changes as agents are added. Analyse
- Compare coordination approaches for a given deployment constraint. Evaluate
1 · Match Approaches to Applications
Section titled “1 · Match Approaches to Applications”Drag each block into a slot, or tap one and then tap a slot.
Each agent learns using only its own experience.
Global information is available during training, but agents act from local observations at deployment.
A value function uses global information during training while the actor stays decentralized.
Individual action values are added to produce a team value.
Individual values are combined by a monotonic mixing function.
2 · Calculate the Joint-Action Space
Section titled “2 · Calculate the Joint-Action Space”Four agents each have five possible actions. How many joint actions are possible?
Every choice by agent 1 combines with every choice by agent 2, and so on. The count is a product, not a sum.
Now six agents, still five actions each.
. Two more agents multiplied the previous answer by 25.
3 · Calculate a VDN Team Value
Section titled “3 · Calculate a VDN Team Value”Under VDN the team value is the sum of the individual values, .
Given , and , calculate .
Agent 2 switches to an action with . What is the new team value?
Only one term changed. You do not need to add all three again.
4 · Explore Joint-Action Growth
Section titled “4 · Explore Joint-Action Growth”Each agent has five actions. Drag the slider to add agents.
At ten agents, how many joint actions are there?
5 · Evaluate the CTDE Trade-off
Section titled “5 · Evaluate the CTDE Trade-off”Ten warehouse robots must act independently at deployment. During training, the simulator provides the complete warehouse state.
Would you prefer fully independent learning or CTDE here? What is the main trade-off in your choice?
AnswersReveal
Independent learning · CTDE · centralized critic · VDN · QMIX.
If one tripped you up it is probably CTDE against centralized critic: CTDE is the paradigm that permits global information during training, and a centralized critic is one way of using it.
.
. Two extra agents multiply by , not by 2.
.
. Under a plain sum, one agent changing action moves the team value by exactly its own change.
. The tenth agent added more joint actions than the first nine combined.
No single right answer. A good one names what it gives up.
CTDE is the strong default: the simulator’s state costs nothing during training, execution is local either way, so refusing it buys nothing. Price: one more component to train and debug.
Independent learning is defensible under time pressure. It needs no new algorithm, keeps per-agent cost flat as the fleet grows, and is the baseline CTDE has to beat. Price: non-stationarity and high variance across seeds.
Review with Flashcards
Section titled “Review with Flashcards”Revisit the 10 key concepts, equations, and intuitions from this chapter.