Skip to content
MARL in Cooperative Environments
Edit this page

1.8Value Decomposition

5 min read

Value decomposition represents a centralized team value through per-agent utilities that can guide decentralized action selection.

In this section you will

  • Write the centralized joint-action value
  • Introduce per-agent utilities and say what they are not
  • State the Individual-Global-Max condition and why it matters
  • Connect the factorization back to credit assignment

Suppose we learn a value for the whole team.

Figure 1
Qtot(h, z, at)\tone{reward}{Q_{\text{tot}}\bigl(\tone{observe}{h},\ \tone{comm}{z},\ \tone{action}{\jointact}\bigr)}
QtotQ_{\text{tot}}
the team value: how good this joint action is here
hh
the agents’ histories
zz
centralized information, available at training time
at\jointact
the joint action being evaluated
One value for what the team did together, exactly the form that can tell combinations apart.

This is a good object to have. It is the joint-action value from Coordination, so it can distinguish (collect, prepare) from (collect, collect) rather than averaging them into one uninformative number.

The trouble is using it to act.

Figure 2
at∗=arg⁡max⁡at  Qtot(h,z,at)\jointact^{*} = \arg\max_{\tone{action}{\jointact}} \; \tone{reward}{Q_{\text{tot}}\bigl(h, z, \tone{action}{\jointact}\bigr)}
arg⁡max⁡at\arg\max_{\jointact}
search over EVERY joint action, the exponentially large set
Choosing the best joint action means searching a space that multiplies with every agent added.

Two things are wrong with that expression, and only one of them is about cost.

It is expensive. The search is over mnm^\nag entries. This is the wall from Centralized and Decentralized Learning, arriving again.

Nobody can evaluate it. Even given infinite compute, the arg⁡max⁡\arg\max needs zz and returns a joint action, so it describes a central controller. At execution time there is no such thing. Each agent has its own observation and must produce its own action.

This is precisely the motivation the textbook gives for value decomposition: centralized values are useful, but decentralized and efficient action selection remains difficult as the joint action space grows.

So we want per-agent values, which each agent can maximise by itself:

Figure 3
Qi(hi, ati)\tone{policy}{Q_\ag\bigl(\tone{observe}{h^\ag},\ \tone{action}{\act{\ag}}\bigr)}
QiQ_\ag
a value over agent i’s own actions, computable from its own history
Something each agent can hold, and maximise, entirely on its own.

Each agent takes arg⁡max⁡\arg\max over its own mm actions. Cheap, local, deployable. But nothing so far connects these to the team value, and without that connection, agents maximising their own QiQ_\ag have no reason to end up anywhere good together.

That connection is the whole idea.

Written out:

Figure 4
arg⁡max⁡atQtot  =  (  arg⁡max⁡at1Q1, …, arg⁡max⁡atnQn  )\arg\max_{\tone{action}{\jointact}} \tone{reward}{Q_{\text{tot}}} \;=\; \Bigl(\; \arg\max_{\act{1}} \tone{policy}{Q_1},\ \dots,\ \arg\max_{\act{\nag}} \tone{policy}{Q_\nag} \;\Bigr)
arg⁡max⁡atQtot\arg\max_{\jointact} Q_{\text{tot}}
the best joint action according to the team value
arg⁡max⁡atiQi\arg\max_{\act{\ag}} Q_\ag
each agent’s own greedy choice, made locally and independently
Everybody choosing greedily and alone lands exactly on the team's best joint action.

Read the equality carefully, because it is doing something remarkable. The left side requires a search over the exponential joint space using centralized information. The right side is n\nag independent searches over mm actions each, using only local information. IGM says these give the same answer.

If you can arrange for that, you get the best of both moments: the team value guides learning, and action selection is local and cheap.

Worth pausing on, because it is easy to miss.

Learning a factored QtotQ_{\text{tot}} means learning what each QiQ_\ag has to be for the factorisation to explain the team’s returns. An agent whose actions consistently coincide with good outcomes acquires a QiQ_\ag that reflects it.

So value decomposition does not merely make action selection tractable. It distributes a signal derived from the shared team reward back to individual agents, which is the credit assignment problem, approached from the other end. That is why the two sections sit next to each other, and why value decomposition belongs to the common-reward setting specifically: without one shared return, there is nothing to decompose.

Knowledge check

Why is the IGM property the thing that makes value decomposition work, rather than just a technical detail?

Select one answer.

  • A centralized team value Qtot(h,z,at)Q_{\text{tot}}(h, z, \jointact) can tell combinations apart, but selecting from it needs an exponential search and centralized information, so no agent can act on it.
  • We want per-agent values Qi(hi,ati)Q_\ag(h^\ag, \act{\ag}) that each agent can maximise locally.
  • The Individual-Global-Max (IGM) property is the required link: each agent’s own greedy action, taken together, equals the team value’s best joint action.
  • IGM constrains the arg max only. Individual values are utilities, not shares of the return.
  • Value decomposition is also a credit assignment mechanism, which is why it belongs specifically to the common-reward setting.