1.9VDN and QMIX
VDN and QMIX are two value-decomposition methods that enforce compatibility between local greedy actions and the team’s preferred joint action.
In this section you will
- Derive VDN’s additive factorization
- Identify the agent interactions additivity cannot represent
- Describe QMIX’s learned monotonic mixing
- State the constraint monotonic mixing still imposes
- Choose between the two factorizations for a given task
VDN: Additive Factorization
Section titled “VDN: Additive Factorization”Value Decomposition Networks take the simplest option available. Assume the team value is the sum of the individual values.
- the team value
- agent i’s individual value, over its own history and action
Nothing is learned about how to combine, the combination is fixed in advance. Only the individual are learned, and they are trained so that their sum explains the team’s returns.
Addition satisfies IGM immediately. Raising any raises , and no agent’s choice can change another agent’s term, so each agent maximising its own value maximises the sum.
Limits of Additive Factorization
Section titled “Limits of Additive Factorization”Addition is a strong assumption, and it is easy to see what it rules out. Take the joint values from Credit Assignment:
| Agent 1 | Agent 2 | |
|---|---|---|
| collect | prepare | 8.0 |
| collect | collect | 2.5 |
| prepare | prepare | 2.0 |
| prepare | collect | 7.5 |
If , then adding the two diagonals of that table must give the same total either way, every term appears exactly once in each pairing:
- the sum for one pairing of local actions
- the sum for the crossed pairing
They do not match, so no choice of and reproduces this table by addition. And the table is not exotic: it just says the agents need to do different things. Addition cannot express “these actions are good together and bad apart”, because it has no term in which the two agents interact.
QMIX: Monotonic Mixing
Section titled “QMIX: Monotonic Mixing”QMIX replaces the fixed sum with a learned function.
- a learned mixing function, constrained to be monotonic in every input
- the same individual values as before, still local
The mixing function is not free. It is constrained to be monotonic in each individual value:
- how the team value responds to raising one agent’s value: never negatively
That constraint is doing something specific, and it is worth being precise about why it is there. Monotonicity is what preserves IGM. If raising can never lower , then each agent taking its own still lands on the joint action the team value ranks highest, so decentralized greedy action selection remains consistent with the centralized value, exactly as under addition.
The architecture that enforces monotonicity is a small network with non-negative weights, conditioned on the state. That detail matters for implementing QMIX and not for understanding it: the idea is a learned combination that is not allowed to go downhill in any agent’s value.
Limits of Monotonic Factorization
Section titled “Limits of Monotonic Factorization”Worth knowing, so the progression does not read as solved.
Consider two agents that must choose differently, the coordination game from Coordination:
| Agent 1 | Agent 2 | |
|---|---|---|
| left | right | 5 |
| right | left | 5 |
| left | left | 0 |
| right | right | 0 |
No monotonic mixing can handle this. Monotonicity forces each agent to have a single greedy action, so their greedy choices combine into either (left, left) or (right, right), both worth 0. The two good joint actions are unreachable by local greedy selection, whatever the mixer.
Selecting a Factorization
Section titled “Selecting a Factorization”Pedagogically, yes.
VDN makes the decomposition idea immediately graspable: one team value, split by addition, agents act locally. Nothing about it is mysterious, and it establishes that the trick is possible.
QMIX shows why the simple version is not enough, and what the minimum generalisation is. Together they trace the actual shape of the argument: decompose the value, discover addition is too rigid, keep only the property you actually need, monotonicity, because monotonicity is what buys IGM.
| VDN | QMIX | |
|---|---|---|
| combination | fixed sum | learned function |
| constraint | additivity | monotonicity in each |
| satisfies IGM | yes | yes |
| representable class | additive team values | monotonic team values |
| what is learned | the | the and the mixer |
Knowledge check
Correct.
Not quite.
VDN fixes the combination as a sum; QMIX learns the combination, restricted to functions monotonic in each individual value so local greedy selection still matches the team value.
That is the distinction. Addition is one monotonic function, so QMIX’s representable class strictly contains VDN’s, and monotonicity is retained precisely because it is what preserves IGM.
QMIX is a better-tuned version of VDN that performs more strongly on benchmarks.
This is the common misreading. The difference is structural rather than a matter of tuning: additive versus monotonic decomposition. Describing it as "better VDN" hides the property that actually changed.
QMIX drops the IGM requirement in exchange for a more expressive value function.
The opposite. QMIX keeps IGM. That is exactly what the monotonicity constraint is protecting. Dropping IGM would mean local greedy choices could disagree with the team value, losing decentralized action selection.
QMIX uses the global state at execution time, while VDN does not.
Neither uses the global state at execution. QMIX’s mixing network is conditioned on the state during training and is not needed to act: each agent selects greedily from its own Q. Both are CTDE methods.
Explanation
Being able to say precisely what changed, the constraint, not the quality, is what separates understanding these methods from reciting them.
VDN and QMIX Summary
Section titled “VDN and QMIX Summary”- VDN assumes . Addition satisfies IGM automatically, and shows that decomposition is possible at all.
- Addition cannot represent interaction. If the diagonals of a joint-value table do not sum equally, no additive decomposition exists.
- QMIX learns a mixing function constrained to be monotonic in every .
- Monotonicity is kept because it is what preserves IGM: local greedy selection still agrees with the centralized value.
- QMIX is a more general form of decomposition, not a tuned VDN. Addition is one monotonic function among many.
- Monotonic decomposition still cannot represent problems requiring agents to choose differently with no way to break the tie.