Skip to content
MARL in Cooperative Environments
Edit this page

1.9VDN and QMIX

6 min read

VDN and QMIX are two value-decomposition methods that enforce compatibility between local greedy actions and the team’s preferred joint action.

In this section you will

  • Derive VDN’s additive factorization
  • Identify the agent interactions additivity cannot represent
  • Describe QMIX’s learned monotonic mixing
  • State the constraint monotonic mixing still imposes
  • Choose between the two factorizations for a given task

Value Decomposition Networks take the simplest option available. Assume the team value is the sum of the individual values.

Figure 1
Qtot=∑i=1nQi(hi,ati)\tone{reward}{Q_{\text{tot}}} = \sum_{\ag=1}^{\nag} \tone{policy}{Q_\ag\bigl(h^\ag, \act{\ag}\bigr)}
QtotQ_{\text{tot}}
the team value
QiQ_\ag
agent i’s individual value, over its own history and action
One team value, assumed to be the sum of the parts.

Nothing is learned about how to combine, the combination is fixed in advance. Only the individual QiQ_\ag are learned, and they are trained so that their sum explains the team’s returns.

Addition satisfies IGM immediately. Raising any QiQ_\ag raises QtotQ_{\text{tot}}, and no agent’s choice can change another agent’s term, so each agent maximising its own value maximises the sum.

Addition is a strong assumption, and it is easy to see what it rules out. Take the joint values from Credit Assignment:

Agent 1Agent 2QtotQ_{\text{tot}}
collectprepare8.0
collectcollect2.5
prepareprepare2.0
preparecollect7.5

If Qtot=Q1+Q2Q_{\text{tot}} = Q_1 + Q_2, then adding the two diagonals of that table must give the same total either way, every term appears exactly once in each pairing:

Figure 2
8.0+7.5⏟= 15.5  =?  2.5+2.0⏟= 4.5\ubrace{reward}{8.0 + 7.5}{= 15.5} \;\overset{?}{=}\; \ubrace{conflict}{2.5 + 2.0}{= 4.5}
8.0+7.58.0 + 7.5
the sum for one pairing of local actions
2.5+2.02.5 + 2.0
the sum for the crossed pairing
Unequal diagonal sums prove that this joint value is not additively separable.

They do not match, so no choice of Q1Q_1 and Q2Q_2 reproduces this table by addition. And the table is not exotic: it just says the agents need to do different things. Addition cannot express “these actions are good together and bad apart”, because it has no term in which the two agents interact.

QMIX replaces the fixed sum with a learned function.

Figure 3
Qtot=fmix(Q1,…,Qn)\tone{reward}{Q_{\text{tot}}} = \tone{comm}{f_{\text{mix}}}\bigl(\tone{policy}{Q_1}, \dots, \tone{policy}{Q_\nag}\bigr)
fmixf_{\text{mix}}
a learned mixing function, constrained to be monotonic in every input
QiQ_\ag
the same individual values as before, still local
How the individual values combine is now learned rather than assumed.

The mixing function is not free. It is constrained to be monotonic in each individual value:

Figure 4
∂Qtot∂Qi  ≥  0for every agent i\frac{\partial \tone{reward}{Q_{\text{tot}}}}{\partial \tone{policy}{Q_\ag}} \;\geq\; 0 \qquad \text{for every agent } \ag
∂Qtot/∂Qi\partial Q_{\text{tot}} / \partial Q_\ag
how the team value responds to raising one agent’s value: never negatively
Whatever the mixing function does, raising one agent's value can never lower the team's.

That constraint is doing something specific, and it is worth being precise about why it is there. Monotonicity is what preserves IGM. If raising QiQ_\ag can never lower QtotQ_{\text{tot}}, then each agent taking its own arg⁡max⁡\arg\max still lands on the joint action the team value ranks highest, so decentralized greedy action selection remains consistent with the centralized value, exactly as under addition.

The architecture that enforces monotonicity is a small network with non-negative weights, conditioned on the state. That detail matters for implementing QMIX and not for understanding it: the idea is a learned combination that is not allowed to go downhill in any agent’s value.

Worth knowing, so the progression does not read as solved.

Consider two agents that must choose differently, the coordination game from Coordination:

Agent 1Agent 2QtotQ_{\text{tot}}
leftright5
rightleft5
leftleft0
rightright0

No monotonic mixing can handle this. Monotonicity forces each agent to have a single greedy action, so their greedy choices combine into either (left, left) or (right, right), both worth 0. The two good joint actions are unreachable by local greedy selection, whatever the mixer.

Pedagogically, yes.

VDN makes the decomposition idea immediately graspable: one team value, split by addition, agents act locally. Nothing about it is mysterious, and it establishes that the trick is possible.

QMIX shows why the simple version is not enough, and what the minimum generalisation is. Together they trace the actual shape of the argument: decompose the value, discover addition is too rigid, keep only the property you actually need, monotonicity, because monotonicity is what buys IGM.

VDNQMIX
combinationfixed sumlearned function
constraintadditivitymonotonicity in each QiQ_\ag
satisfies IGMyesyes
representable classadditive team valuesmonotonic team values
what is learnedthe QiQ_\agthe QiQ_\ag and the mixer

Knowledge check

What is the most accurate statement of how QMIX differs from VDN?

Select one answer.

  • VDN assumes Qtot=∑iQiQ_{\text{tot}} = \sum_\ag Q_\ag. Addition satisfies IGM automatically, and shows that decomposition is possible at all.
  • Addition cannot represent interaction. If the diagonals of a joint-value table do not sum equally, no additive decomposition exists.
  • QMIX learns a mixing function fmix(Q1,…,Qn)f_{\text{mix}}(Q_1, \dots, Q_\nag) constrained to be monotonic in every QiQ_\ag.
  • Monotonicity is kept because it is what preserves IGM: local greedy selection still agrees with the centralized value.
  • QMIX is a more general form of decomposition, not a tuned VDN. Addition is one monotonic function among many.
  • Monotonic decomposition still cannot represent problems requiring agents to choose differently with no way to break the tie.