1.2Coordination
Coordination is the problem of learning individual actions whose combination advances a shared objective.
In this section you will
- Explain why aligned objectives do not imply compatible actions
- Read joint-action dependence off a small payoff table
- Define a joint-action value and say what it is a function of
Aligned Objectives and Interdependent Actions
Section titled “Aligned Objectives and Interdependent Actions”Two agents in a kitchen. When an order is served, both receive the same payment. Neither can earn anything alone.
The simplest coordination problem there is. Same reward, same goal, separate decisions.
Formally, this is a common-reward game: every agent has the same reward function, so the team’s interests are identical by construction.
- agent i’s reward function
- the single function they all share
You would think that settles it. If everybody wants the same thing, and nothing anybody does can help themselves at another’s expense, what is left to go wrong?
Almost everything.
The textbook version of this makes it unmissable. Albrecht, Christianos and Schäfer use a coordination game in which two agents receive a positive reward only when their actions agree, and nothing tells either agent which of the matching options to pick. Both agents want to agree. Wanting it does not make it happen.
Joint-Action Dependence
Section titled “Joint-Action Dependence”Back to the kitchen, and to what actually arrives at the environment.
- the joint action: what the team did this step
- one agent’s contribution to it, chosen from local information
Four possibilities at one step, with the same two agents and the same shared reward:
| Agent 1 | Agent 2 | Team outcome |
|---|---|---|
| get ingredient | prepare station | both halves of the next step are ready |
| get ingredient | get ingredient | two trips for one ingredient |
| prepare station | prepare station | the station is prepared twice |
| wait | wait | nothing happens |
Every one of those rows involves two agents who want an order served. Rows two and three are agents doing genuinely useful things, preparing a station is not a mistake, that happen to be the same useful thing at the same moment. Row four is two agents each sensibly waiting for the other to commit.
None of this is bad behaviour. It is uncoordinated behaviour, which is a different thing and the whole subject of this chapter.
Joint-Action Values
Section titled “Joint-Action Values”Now the formal consequence, and it is the piece worth carrying forward.
In single-agent reinforcement learning you can ask how good an action is in a state, and write the answer down as . Try that here.
- an individual action value: how good agent 1’s action is, with agent 2’s action left out
- a joint-action value: how good this combination is
The left-hand form cannot represent the kitchen table above. “Get ingredient” appears in row one, where it is excellent, and row two, where it is wasted. Any single number attached to “get ingredient” has to be some average across what agent 2 might do, and that average is not the value of anything the agent can actually bring about.
The right-hand form has room for the distinction. Give it both actions and it can say that (get, prepare) is worth a lot and (get, get) is worth little.
And the cost is real: has an entry for every joint action, so it grows exponentially in the number of agents, and, worse for our purposes, an agent choosing from it would need to know what everybody else is about to do. Centralized and Decentralized Learning takes that problem seriously.
Knowledge check
Correct.
Not quite.
There is not enough information. It depends on what agent 2 does at the same step.
Correct, and getting comfortable with this answer is the point of the section. The action is excellent if agent 2 prepares the station and wasteful if agent 2 also goes for the ingredient. Nothing in the state, the reward, or the action itself settles it.
Yes. The order needs an ingredient and none has been collected.
Tempting, and it is the reasoning that works in single-agent RL. But the same reasoning is available to agent 2, and if both act on it the team makes two trips for one ingredient. An action that is obviously right for everyone is exactly how duplication happens.
No. Agent 1 should wait to see what agent 2 does first.
They act simultaneously, so there is no "first", waiting to see is not an available strategy. And if both agents wait, nothing happens at all, which is row four of the table.
Yes, as long as agent 1’s policy is well trained.
Training quality is not what is missing. Even an optimal agent 1 cannot make "get ingredient" good in isolation, because its value is a property of the combination rather than of the action.
Explanation
“It depends on what the other agent does” is the correct answer far more often than it feels like it should be.
Coordination Summary
Section titled “Coordination Summary”- A common-reward game gives every agent the same reward function. Their objectives are identical by construction.
- Identical objectives do not produce coordinated actions. Agents can want the same outcome and still duplicate work, interfere, or both wait.
- Poor coordination is usually not bad behaviour. It is two sensible choices that do not fit together.
- An individual action value must average over what partners do, hiding the effect that decides the outcome.
- A joint-action value can distinguish combinations, at a cost in size, and in information no single agent has.