Skip to content
MARL in Cooperative Environments
Edit this page

1.2Coordination

5 min read

Coordination is the problem of learning individual actions whose combination advances a shared objective.

In this section you will

  • Explain why aligned objectives do not imply compatible actions
  • Read joint-action dependence off a small payoff table
  • Define a joint-action value and say what it is a function of

Aligned Objectives and Interdependent Actions

Section titled “Aligned Objectives and Interdependent Actions”

Two agents in a kitchen. When an order is served, both receive the same payment. Neither can earn anything alone.

Two robot agents working in one kitchen. The left agent is at a chopping board, the right agent at a stove with a pot. They share one counter. A single order ticket above the serving hatch is the reward, and it is paid to both of them together only when the order is complete. Agent 1Agent 2ONE ORDERone shared rewardneither is paid alone

The simplest coordination problem there is. Same reward, same goal, separate decisions.

Formally, this is a common-reward game: every agent has the same reward function, so the team’s interests are identical by construction.

Figure 1
R1=R2=⋯=Rn=R\tone{reward}{R_1} = \tone{reward}{R_2} = \cdots = \tone{reward}{R_\nag} = \tone{reward}{R}
RiR_\ag
agent i’s reward function
RR
the single function they all share
Identical reward functions. There is nothing left for the agents to disagree about.

You would think that settles it. If everybody wants the same thing, and nothing anybody does can help themselves at another’s expense, what is left to go wrong?

Almost everything.

The textbook version of this makes it unmissable. Albrecht, Christianos and Schäfer use a coordination game in which two agents receive a positive reward only when their actions agree, and nothing tells either agent which of the matching options to pick. Both agents want to agree. Wanting it does not make it happen.

Back to the kitchen, and to what actually arrives at the environment.

Figure 2
at=(at1, …, atn)\tone{action}{\jointact} = \bigl(\tone{policy}{\act{1}},\ \dots,\ \tone{policy}{\act{\nag}}\bigr)
at\jointact
the joint action: what the team did this step
ati\act{i}
one agent’s contribution to it, chosen from local information
The environment sees the combination, never the parts.

Four possibilities at one step, with the same two agents and the same shared reward:

Agent 1Agent 2Team outcome
get ingredientprepare stationboth halves of the next step are ready
get ingredientget ingredienttwo trips for one ingredient
prepare stationprepare stationthe station is prepared twice
waitwaitnothing happens

Every one of those rows involves two agents who want an order served. Rows two and three are agents doing genuinely useful things, preparing a station is not a mistake, that happen to be the same useful thing at the same moment. Row four is two agents each sensibly waiting for the other to commit.

None of this is bad behaviour. It is uncoordinated behaviour, which is a different thing and the whole subject of this chapter.

Now the formal consequence, and it is the piece worth carrying forward.

In single-agent reinforcement learning you can ask how good an action is in a state, and write the answer down as Q(s,a)Q(s, a). Try that here.

Figure 3
Q(s,a1)versusQ(s,a1,a2)\dbox{conflict}{Q(\tone{observe}{s}, \tone{action}{a^{1}})} \qquad\text{versus}\qquad \cbox{action}{Q(\tone{observe}{s}, \tone{action}{a^{1}}, \tone{action}{a^{2}})}
Q(s,a1)Q(s, a^{1})
an individual action value: how good agent 1’s action is, with agent 2’s action left out
Q(s,a1,a2)Q(s, a^{1}, a^{2})
a joint-action value: how good this combination is
The dashed form has to average away the very thing that decides the outcome. The solid form can tell the combinations apart.

The left-hand form cannot represent the kitchen table above. “Get ingredient” appears in row one, where it is excellent, and row two, where it is wasted. Any single number attached to “get ingredient” has to be some average across what agent 2 might do, and that average is not the value of anything the agent can actually bring about.

The right-hand form has room for the distinction. Give it both actions and it can say that (get, prepare) is worth a lot and (get, get) is worth little.

And the cost is real: Q(s,at)Q(s, \jointact) has an entry for every joint action, so it grows exponentially in the number of agents, and, worse for our purposes, an agent choosing from it would need to know what everybody else is about to do. Centralized and Decentralized Learning takes that problem seriously.

Knowledge check

An order has just arrived and the ingredient store is full. Is 'get ingredient' a good action for agent 1?

Select one answer.

  • A common-reward game gives every agent the same reward function. Their objectives are identical by construction.
  • Identical objectives do not produce coordinated actions. Agents can want the same outcome and still duplicate work, interfere, or both wait.
  • Poor coordination is usually not bad behaviour. It is two sensible choices that do not fit together.
  • An individual action value Q(s,ai)Q(s, a^\ag) must average over what partners do, hiding the effect that decides the outcome.
  • A joint-action value Q(s,at)Q(s, \jointact) can distinguish combinations, at a cost in size, and in information no single agent has.