Skip to content
MARL in Cooperative Environments
Edit this page

1.7Credit Assignment

5 min read

Credit assignment asks how a shared outcome should influence the agents and actions that produced it.

In this section you will

  • Separate temporal credit assignment from multi-agent credit assignment
  • Use joint-action values to supply the missing context
  • Construct a counterfactual comparison for one agent’s action

Temporal and Multi-Agent Credit Assignment

Section titled “Temporal and Multi-Agent Credit Assignment”

“Credit assignment” refers to two different difficulties, and separating them helps.

Temporal credit assignment asks which of my earlier actions mattered. Agent 1 served the order at step 5, but the collect at step 1 was just as necessary. This exists in single-agent reinforcement learning too, and discounted returns and value functions are the standard machinery for it.

Multi-agent credit assignment asks which agent mattered. Both agents received +10+10. Did agent 2’s step-2 wait help, or was it idling? Would the order still have gone out if it had done something else? This one is new, and it is what this chapter is about.

Note the shape of the difficulty. It is not that the information is noisy. It is that the reward was never a sum of per-agent contributions in the first place, nothing was added up, so nothing can be taken apart by inspecting the total.

Coordination introduced the two forms. Bring them back, because this is what they are for.

Figure 1
Q(s,at)\cbox{action}{Q\bigl(\tone{observe}{s}, \tone{action}{\jointact}\bigr)}
Q(s,at)Q(s, \jointact)
the value of this state together with this particular combination of actions
A value that can hold different numbers for different combinations of who did what.

Because this function is defined on the whole joint action, it can assign different values to different combinations, and that is more information about how the agents contributed than a single team reward carries. The textbook identifies joint-action values as one approach to disentangling contributions for exactly this reason.

Consider a two-agent step, with values a joint-action function could hold:

Agent 1Agent 2Q(s,at)Q(s, \jointact)
collectprepare8.0
collectcollect2.5
prepareprepare2.0
preparecollect7.5

Read down the first column. When agent 1 collects, the outcome is 8.0 or 2.5 depending entirely on agent 2. A per-agent value Q1(s,a1)Q_1(s, a^1) would have to report one number for “collect”, somewhere in the middle, describing neither case.

The table also suggests how to extract an individual contribution from a joint value, and the idea is simple enough to state in a sentence.

What would have happened if agent 1 had chosen a different action, while everybody else behaved the same way?

That question is answerable from the table. Holding agent 2 at “prepare”, agent 1 collecting is worth 8.0 and preparing is worth 2.0. The difference is attributable to agent 1, because nothing else changed.

Figure 2
contribution of i  ≈  Q(s,at)  −  ∑a′πi(a′) Q(s,(a′,at−i))⏟the same step with agent i’s action replaced\text{contribution of } \ag \;\approx\; \tone{action}{Q(s, \jointact)} \;-\; \ubrace{observe}{\textstyle\sum_{a'} \pol{\ag}(a') \, Q\bigl(s, (a', \jointact^{-\ag})\bigr)}{the same step with agent i's action replaced}
at−i\jointact^{-\ag}
what every OTHER agent did, held fixed
a′a'
an alternative action agent i could have taken instead
πi(a′)\pol{\ag}(a')
how likely agent i was to take that alternative
Compare what happened against the average of what would have happened if only this agent had acted differently.

Hold the partners fixed, vary one agent, and the change in value is that agent’s contribution to this step. This is counterfactual credit assignment, and it is a genuinely satisfying answer to “what did I do?”

Knowledge check

In the kitchen trajectory above, agent 2 'waits' at step 2 and the episode ends with +10. What can a learner conclude about that wait?

Select one answer.

  • Temporal credit assignment asks which of my earlier actions mattered; multi-agent credit assignment asks which agent mattered. This chapter is about the second.
  • A common reward makes it hard by construction: the same number goes to every agent, and it was never a sum of per-agent contributions.
  • A joint-action value Q(s,at)Q(s, \jointact) can distinguish combinations, and so carries information about contributions that the team reward does not.
  • Counterfactual reasoning extracts an individual contribution: compare what happened against what would have happened if only that agent had acted differently, holding partners fixed.
  • Handing out per-agent rewards would make credit legible and would stop the problem being cooperative.