1.8Value Decomposition
Value decomposition represents a centralized team value through per-agent utilities that can guide decentralized action selection.
In this section you will
- Write the centralized joint-action value
- Introduce per-agent utilities and say what they are not
- State the Individual-Global-Max condition and why it matters
- Connect the factorization back to credit assignment
Centralized Team Value
Section titled “Centralized Team Value”Suppose we learn a value for the whole team.
- the team value: how good this joint action is here
- the agents’ histories
- centralized information, available at training time
- the joint action being evaluated
This is a good object to have. It is the joint-action value from Coordination, so it can distinguish (collect, prepare) from (collect, collect) rather than averaging them into one uninformative number.
The trouble is using it to act.
- search over EVERY joint action, the exponentially large set
Two things are wrong with that expression, and only one of them is about cost.
It is expensive. The search is over entries. This is the wall from Centralized and Decentralized Learning, arriving again.
Nobody can evaluate it. Even given infinite compute, the needs and returns a joint action, so it describes a central controller. At execution time there is no such thing. Each agent has its own observation and must produce its own action.
This is precisely the motivation the textbook gives for value decomposition: centralized values are useful, but decentralized and efficient action selection remains difficult as the joint action space grows.
Individual Utilities
Section titled “Individual Utilities”So we want per-agent values, which each agent can maximise by itself:
- a value over agent i’s own actions, computable from its own history
Each agent takes over its own actions. Cheap, local, deployable. But nothing so far connects these to the team value, and without that connection, agents maximising their own have no reason to end up anywhere good together.
That connection is the whole idea.
Individual-Global-Max
Section titled “Individual-Global-Max”Written out:
- the best joint action according to the team value
- each agent’s own greedy choice, made locally and independently
Read the equality carefully, because it is doing something remarkable. The left side requires a search over the exponential joint space using centralized information. The right side is independent searches over actions each, using only local information. IGM says these give the same answer.
If you can arrange for that, you get the best of both moments: the team value guides learning, and action selection is local and cheap.
Value Decomposition and Credit Assignment
Section titled “Value Decomposition and Credit Assignment”Worth pausing on, because it is easy to miss.
Learning a factored means learning what each has to be for the factorisation to explain the team’s returns. An agent whose actions consistently coincide with good outcomes acquires a that reflects it.
So value decomposition does not merely make action selection tractable. It distributes a signal derived from the shared team reward back to individual agents, which is the credit assignment problem, approached from the other end. That is why the two sections sit next to each other, and why value decomposition belongs to the common-reward setting specifically: without one shared return, there is nothing to decompose.
Knowledge check
Correct.
Not quite.
Without it, agents maximising their own values could jointly select an action the team value considers poor.
Exactly. Per-agent values are only useful if acting greedily on them produces a good joint action. IGM is precisely the guarantee that local greedy selection and centralized greedy selection agree, which is what lets you train centrally and act locally.
It guarantees the individual values sum to the team value.
It does not. IGM constrains only which joint action comes out on top. Summation is one particular way to satisfy it, VDN’s way, not what the property says.
It makes the joint action space smaller.
The joint action space is unchanged; it is still m to the power n. What changes is that nobody has to search it, each agent searches its own m actions instead.
It ensures each agent receives the reward it earned.
Appealing, and stronger than what IGM claims. The individual values are utilities whose ordering yields the right joint choice, not accounting statements about earned reward.
Explanation
IGM is the bridge between the two moments. Everything in the next section is a way of building it.
Value Decomposition Summary
Section titled “Value Decomposition Summary”- A centralized team value can tell combinations apart, but selecting from it needs an exponential search and centralized information, so no agent can act on it.
- We want per-agent values that each agent can maximise locally.
- The Individual-Global-Max (IGM) property is the required link: each agent’s own greedy action, taken together, equals the team value’s best joint action.
- IGM constrains the arg max only. Individual values are utilities, not shares of the return.
- Value decomposition is also a credit assignment mechanism, which is why it belongs specifically to the common-reward setting.