Skip to content
MARL in Cooperative Environments
Edit this page

1.12Learning Coordinated Behaviour

3 min read

What you will be able to do

  • Drive a joint action by hand and see that neither agent can advance the task alone.
  • Watch the joint action space grow as agents are added.
  • Train independent learners and value decomposition on the same task.
  • Compare them on success rate, steps and behaviour, and say what actually differs.

Everything here runs in your browser. Nothing is installed and nothing is sent anywhere.

  • Two agents must hold both switches down on the same step.
  • Each agent chooses from five actions: up, down, left, right, stay.
  • The team receives the reward together, or not at all.

A switch held alone earns nothing. That is the whole design.

Pick an action for each agent, then take the joint action. The environment cannot advance on one agent’s decision.

Agent A
Agent B
Step
0
Reward
0.00
Return
0.00

Try to reach both switches on the same step. Then try arriving early with one agent and see what the reward does while you wait.

Each agent chooses privately from its own five actions, and the environment responds to the combination.

∣A∣=∏i=1n∣Ai∣|\mathcal{A}| = \prod_{i=1}^{n} |\mathcal{A}_i|

Idea: Adding agents quickly increases the number of possible joint decisions.

Drag the slider. Five actions per agent throughout.

The bars are on a linear scale on purpose. A log axis would turn exponential growth into a straight line, which is the opposite of the point.

Two methods, and the difference between them is one line of arithmetic.

Qi(oi,ai)←Qi(oi,ai)+α[r+γmax⁡ai′Qi(oi′,ai′)−Qi(oi,ai)]Q_i(o_i,a_i) \leftarrow Q_i(o_i,a_i) + \alpha \left[ r + \gamma \max_{a_i'}Q_i(o_i',a_i') - Q_i(o_i,a_i) \right]

Idea: Each agent credits itself with the whole team reward and has no representation of the other agent.

Qtot=Q1+Q2Q_{\mathrm{tot}} = Q_1 + Q_2

Idea: Individual action-values are combined into a team-level value during learning, so one shared error is applied to both agents.

Under both methods each agent acts on its own observation: its own cell, plus one bit saying whether the other agent is already in position.

Train each method, then replay what it learned on the same grid you just played on.

  • independent
  • VDN
MethodSuccessStepsReturn
Train a method to see its numbers.

What you should see at the default settings:

  • Both methods reach a 100% success rate and solve the task in 4 steps, which is optimal.
  • Independent learning gets there sooner.
  • The replays are indistinguishable: both walk straight to the switches and hold.

The penalty control changes that. Set Penalty for holding a switch alone to −1-1 or −2-2, reset training, and train both again.

Now standing on a switch while your partner is still walking costs something, so arriving early is no longer free and a fixed plan stops being safe. Watch what happens to the success rate and to how reliably each method converges across the run.

What changed during training under VDN, and what remained decentralized when the agents acted?

Answer it before moving on. There is no solution section for this lab, because the answer is visible in what you just ran rather than in code you were asked to write.