Every card shows its question first. Reveal one at a time to test yourself, or
reveal all of them to use the deck as a revision sheet.
10 cards
Intuition
Why does sharing the same objective not automatically produce coordinated behaviour?
Intuition
Agents can choose individually reasonable actions that combine poorly.
A shared goal tells agents what the team wants, but not necessarily how their actions should fit together.
Scenario
Why can the quality of one agent's action depend on the actions chosen by other agents?
Scenario
Because actions can be complementary or conflicting.
FETCH may be useful if another agent chooses COOK, and wasteful if both agents choose FETCH.
Equation
What does this tell us?
∣Ajoint∣=kn
Equation
If n agents each have k possible actions, there are kn possible joint actions.
The number of combinations grows exponentially with the number of agents.
k
actions available to one agent
n
number of agents
Concept
Why does coordination become harder as the number of agents increases?
Concept
The number of possible joint actions grows rapidly, and each agent must reason about how its decisions interact with those of more agents.
Concept
What is independent learning in MARL?
Concept
Each agent learns its own policy while treating the other agents as part of the environment.
Intuition
Why can independent learning create non-stationarity?
Intuition
Other agents are also updating their policies.
When another agent changes its behaviour, the environment experienced by the learner effectively changes too.
Concept
What is Centralized Training with Decentralized Execution?
Concept
CTDE allows agents to use richer, team-level or global information during training, while requiring each agent to act using only locally available information at execution.
Learn together. Act locally.
Concept
What is the credit assignment problem in cooperative MARL?
Concept
It is the problem of determining how individual agents or actions contributed to a shared team outcome.
A team reward may say the team succeeded without explaining who contributed what.
Equation
What does the VDN equation mean?
Qtot=i∑Qi
Equation
VDN represents the team's value as the sum of the individual agent values.
Qi
local agent value
Qtot
team value
Each agent estimates a local value, and these values combine into the total team value.
Distinction
How does QMIX differ from VDN, and what does its monotonicity constraint mean?
Distinction
VDN combines individual values using a simple sum. QMIX uses a learned mixing network to combine them more flexibly.
∂Qi∂Qtot≥0
The constraint means that increasing an agent's local value cannot decrease the estimated team value.