1.7Credit Assignment
Credit assignment asks how a shared outcome should influence the agents and actions that produced it.
In this section you will
- Separate temporal credit assignment from multi-agent credit assignment
- Use joint-action values to supply the missing context
- Construct a counterfactual comparison for one agent’s action
Temporal and Multi-Agent Credit Assignment
Section titled “Temporal and Multi-Agent Credit Assignment”“Credit assignment” refers to two different difficulties, and separating them helps.
Temporal credit assignment asks which of my earlier actions mattered. Agent 1 served the order at step 5, but the collect at step 1 was just as necessary. This exists in single-agent reinforcement learning too, and discounted returns and value functions are the standard machinery for it.
Multi-agent credit assignment asks which agent mattered. Both agents received . Did agent 2’s step-2 wait help, or was it idling? Would the order still have gone out if it had done something else? This one is new, and it is what this chapter is about.
Note the shape of the difficulty. It is not that the information is noisy. It is that the reward was never a sum of per-agent contributions in the first place, nothing was added up, so nothing can be taken apart by inspecting the total.
Credit from Joint-Action Values
Section titled “Credit from Joint-Action Values”Coordination introduced the two forms. Bring them back, because this is what they are for.
- the value of this state together with this particular combination of actions
Because this function is defined on the whole joint action, it can assign different values to different combinations, and that is more information about how the agents contributed than a single team reward carries. The textbook identifies joint-action values as one approach to disentangling contributions for exactly this reason.
Consider a two-agent step, with values a joint-action function could hold:
| Agent 1 | Agent 2 | |
|---|---|---|
| collect | prepare | 8.0 |
| collect | collect | 2.5 |
| prepare | prepare | 2.0 |
| prepare | collect | 7.5 |
Read down the first column. When agent 1 collects, the outcome is 8.0 or 2.5 depending entirely on agent 2. A per-agent value would have to report one number for “collect”, somewhere in the middle, describing neither case.
Counterfactual Credit
Section titled “Counterfactual Credit”The table also suggests how to extract an individual contribution from a joint value, and the idea is simple enough to state in a sentence.
What would have happened if agent 1 had chosen a different action, while everybody else behaved the same way?
That question is answerable from the table. Holding agent 2 at “prepare”, agent 1 collecting is worth 8.0 and preparing is worth 2.0. The difference is attributable to agent 1, because nothing else changed.
- what every OTHER agent did, held fixed
- an alternative action agent i could have taken instead
- how likely agent i was to take that alternative
Hold the partners fixed, vary one agent, and the change in value is that agent’s contribution to this step. This is counterfactual credit assignment, and it is a genuinely satisfying answer to “what did I do?”
Knowledge check
Correct.
Not quite.
Nothing from the team reward alone. Deciding whether the wait helped needs a comparison against what would have happened had agent 2 acted differently.
Correct, and this is the whole motivation for counterfactual reasoning. The wait may have been essential, the ingredient was not chopped yet, or pure idling. The +10 is identical either way, so the comparison has to come from somewhere else.
That waiting was a good action, since the episode succeeded.
This is the trap the shared reward sets. Every action in a successful episode gets the same positive signal, including the useless ones. Learning from that directly reinforces idling whenever the rest of the team happens to succeed.
That waiting contributed one ninth of the reward, since nine actions were taken.
Dividing equally is a real technique and it is also plainly wrong as an account of contribution. It would credit an idle step exactly as much as the serve. The reward never decomposed into per-action shares, so any split is an assumption rather than a reading.
That agent 2 should be given its own reward function instead.
That would make credit legible and would stop the problem being cooperative. You would now have a mixed-incentive game, where agent 2 can score by doing things that do not serve orders. The point of this chapter is recovering individual signal *while keeping* the shared reward.
Explanation
“The same reward is consistent with several different stories” is the sentence to keep.
Credit Assignment Summary
Section titled “Credit Assignment Summary”- Temporal credit assignment asks which of my earlier actions mattered; multi-agent credit assignment asks which agent mattered. This chapter is about the second.
- A common reward makes it hard by construction: the same number goes to every agent, and it was never a sum of per-agent contributions.
- A joint-action value can distinguish combinations, and so carries information about contributions that the team reward does not.
- Counterfactual reasoning extracts an individual contribution: compare what happened against what would have happened if only that agent had acted differently, holding partners fixed.
- Handing out per-agent rewards would make credit legible and would stop the problem being cooperative.