Skip to content
MARL in Cooperative Environments
Edit this page

1.13Coordinate Worksheet

3 min read

This five-part digital activity consolidates the chapter’s central coordination ideas. You will first match learning approaches to their applications, then calculate a joint-action-space size and a VDN team value. An interactive graph shows how joint-action counts grow with the number of agents. The final prompt asks you to weigh the training benefits of CTDE against its additional complexity. Every item except that trade-off response is automatically checkable.

What you will be able to do

  • Match common MARL training approaches to the problems they address. Understand
  • Calculate the size and value of simple joint-action problems. Apply
  • Interpret how coordination difficulty changes as agents are added. Analyse
  • Compare coordination approaches for a given deployment constraint. Evaluate

Drag each block into a slot, or tap one and then tap a slot.

  1. Each agent learns using only its own experience.

  2. Global information is available during training, but agents act from local observations at deployment.

  3. A value function uses global information during training while the actor stays decentralized.

  4. Individual action values are added to produce a team value.

  5. Individual values are combined by a monotonic mixing function.

Four agents each have five possible actions. How many joint actions are possible?

joint actions

Now six agents, still five actions each.

joint actions

Under VDN the team value is the sum of the individual values, Qtot=∑iQiQ_{\text{tot}} = \sum_i Q_i.

Given Q1=3.2Q_1 = 3.2, Q2=2.6Q_2 = 2.6 and Q3=4.1Q_3 = 4.1, calculate QtotQ_{\text{tot}}.

Agent 2 switches to an action with Q2=3.4Q_2 = 3.4. What is the new team value?

Each agent has five actions. Drag the slider to add agents.

Joint actions against Number of agents. 1: 5, 2: 25, 3: 125, 4: 625, 5: 3,125, 6: 15,625, 7: 78,125, 8: 390,625, 9: 1,953,125, 10: 9,765,625 02,441,4064,882,8137,324,2199,765,62512345678910Number of agentsJoint actions

At ten agents, how many joint actions are there?

joint actions

Ten warehouse robots must act independently at deployment. During training, the simulator provides the complete warehouse state.

Would you prefer fully independent learning or CTDE here? What is the main trade-off in your choice?

AnswersReveal
Answer to 1.

Independent learning · CTDE · centralized critic · VDN · QMIX.

If one tripped you up it is probably CTDE against centralized critic: CTDE is the paradigm that permits global information during training, and a centralized critic is one way of using it.

Answer to 2.1.

54=6255^4 = 625.

Answer to 2.2.

56=15,6255^6 = 15{,}625. Two extra agents multiply by 52=255^2 = 25, not by 2.

Answer to 3.1.

3.2+2.6+4.1=9.93.2 + 2.6 + 4.1 = 9.9.

Answer to 3.2.

9.9−2.6+3.4=10.79.9 - 2.6 + 3.4 = 10.7. Under a plain sum, one agent changing action moves the team value by exactly its own change.

Answer to 4.1.

510=9,765,6255^{10} = 9{,}765{,}625. The tenth agent added more joint actions than the first nine combined.

Answer to 5.1.

No single right answer. A good one names what it gives up.

CTDE is the strong default: the simulator’s state costs nothing during training, execution is local either way, so refusing it buys nothing. Price: one more component to train and debug.

Independent learning is defensible under time pressure. It needs no new algorithm, keeps per-agent cost flat as the fleet grows, and is the baseline CTDE has to beat. Price: non-stationarity and high variance across seeds.

Revisit the 10 key concepts, equations, and intuitions from this chapter.

Open Coordinate Flashcards