1.12Learning Coordinated Behaviour
What you will be able to do
- Drive a joint action by hand and see that neither agent can advance the task alone.
- Watch the joint action space grow as agents are added.
- Train independent learners and value decomposition on the same task.
- Compare them on success rate, steps and behaviour, and say what actually differs.
Everything here runs in your browser. Nothing is installed and nothing is sent anywhere.
Mission
Section titled “Mission”- Two agents must hold both switches down on the same step.
- Each agent chooses from five actions: up, down, left, right, stay.
- The team receives the reward together, or not at all.
A switch held alone earns nothing. That is the whole design.
Try it yourself
Section titled “Try it yourself”Pick an action for each agent, then take the joint action. The environment cannot advance on one agent’s decision.
Try to reach both switches on the same step. Then try arriving early with one agent and see what the reward does while you wait.
Joint actions
Section titled “Joint actions”Each agent chooses privately from its own five actions, and the environment responds to the combination.
Key Equation: Joint Action Space
Section titled “Key Equation: Joint Action Space”Idea: Adding agents quickly increases the number of possible joint decisions.
Drag the slider. Five actions per agent throughout.
The bars are on a linear scale on purpose. A log axis would turn exponential growth into a straight line, which is the opposite of the point.
How the agents learn
Section titled “How the agents learn”Two methods, and the difference between them is one line of arithmetic.
Key Equation: Independent Q-Learning
Section titled “Key Equation: Independent Q-Learning”Idea: Each agent credits itself with the whole team reward and has no representation of the other agent.
Key Equation: Value Decomposition
Section titled “Key Equation: Value Decomposition”Idea: Individual action-values are combined into a team-level value during learning, so one shared error is applied to both agents.
Under both methods each agent acts on its own observation: its own cell, plus one bit saying whether the other agent is already in position.
Train and compare
Section titled “Train and compare”Train each method, then replay what it learned on the same grid you just played on.
- independent
- VDN
| Method | Success | Steps | Return |
|---|---|---|---|
| Train a method to see its numbers. | |||
What you should see at the default settings:
- Both methods reach a 100% success rate and solve the task in 4 steps, which is optimal.
- Independent learning gets there sooner.
- The replays are indistinguishable: both walk straight to the switches and hold.
Make the comparison bite
Section titled “Make the comparison bite”The penalty control changes that. Set Penalty for holding a switch alone to or , reset training, and train both again.
Now standing on a switch while your partner is still walking costs something, so arriving early is no longer free and a fixed plan stops being safe. Watch what happens to the success rate and to how reliably each method converges across the run.
One question to take away
Section titled “One question to take away”What changed during training under VDN, and what remained decentralized when the agents acted?
Answer it before moving on. There is no solution section for this lab, because the answer is visible in what you just ran rather than in code you were asked to write.