Skip to content
MARL in Cooperative Environments
Edit this page

6Cooperative Rewards

7 min read

Cooperative agents pursue one shared objective, represented here by a common reward.

In this section you will

  • Contrast per-agent payments with one shared team reward
  • Write the reward as a function of the state and the joint action
  • State the team objective as an expected discounted return
  • Separate aligned objectives from shared credit
  • Read a shaped cooperative reward term by term

There are really only two options, and choosing between them decides what kind of problem you have.

Pay them separately. Agent 1 earns per vegetable chopped. Agent 2 earns per minute the stove is at temperature. Both can now be measured alone, which is convenient, and immediately wrong. A chef who chops a mountain of onions nobody needs is doing beautifully by that measure. Neither number mentions the order going out.

Pay them once, together. One payment arrives when the dish is served, and both of them receive it. Nobody can earn anything by looking productive. The only way to be paid is for the order to be complete.

The second is the cooperative setting.

Two robot agents working in one kitchen. The left agent is at a chopping board, the right agent at a stove with a pot. They share one counter. A single order ticket above the serving hatch is the reward, and it is paid to both of them together only when the order is complete. Agent 1Agent 2ONE ORDERone shared rewardneither is paid alone

The cooperative reward, drawn. One ticket, both agents, paid only when the order is served. Neither agent has a score of its own.

Figure 1
rt=R(st, at)\tone{reward}{\rew} = R\bigl(\tone{observe}{\st},\ \tone{action}{\jointact}\bigr)
rt\rew
the team reward at step t; every agent receives this same value
RR
the reward function, defined on the state and the joint action together
st\st
the state: the situation the environment is in at step t
at\jointact
the joint action, which is why no per-agent reward exists
One number for the whole team, and it depends on what everybody did.

Two features of that line carry the whole section.

It is a function of at\jointact, the joint action, not of any single agent’s choice. There is no RiR^i anywhere in the definition, so there is nothing to evaluate for one agent on its own. Ask “what did chopping earn?” and the question has no answer, because the reward function was never given one agent’s action as an input.

And every agent receives the identical value of rt\rew. In the game-theory vocabulary that makes it a common-reward game, and it is a modelling choice rather than something discovered about the world. You are declaring that this team wins and loses together.

Give every agent the same reward and their objectives are aligned: improving the team return benefits them all. This does not mean their actions will automatically be coordinated. Both kitchen agents want the order served, and they can still both reach for the same pan, or both wait for the other to start. Aligned objectives, interfering actions, the rest of this resource is mostly about that gap.

One serving is not the goal; a good service is. The team is not chasing the next reward but the discounted sum over the whole episode, and because both the environment and the policies can be random, what gets optimised is its expectation.

Figure 2
J(π)=Eπ[  ∑t=0Tγt rt  ]\tone{policy}{J(\boldsymbol{\pi})} = \E_{\boldsymbol{\pi}}\Bigl[\; \cbox{reward}{\sum_{t=0}^{T} \gamma^{t}\, \rew} \;\Bigr]
J(π)J(\boldsymbol{\pi})
the team objective, a function of the whole joint policy
π\boldsymbol{\pi}
the joint policy: the tuple of every agent’s policy
Eπ\E_{\boldsymbol{\pi}}
average over episodes run under that joint policy
∑t=0Tγtrt\sum_{t=0}^{T} \gamma^{t} \rew
the boxed part: the discounted return of one episode, summed to the episode’s last step T
One objective for the whole team, and it depends on every agent's policy at once.

Read the left-hand side carefully, because it is doing something unusual. There is exactly one objective, and its argument is the whole joint policy π\boldsymbol{\pi}.

An agent that improves its own policy while its partners hold still has changed JJ. An agent that improves its own policy while its partners also change may have made JJ worse. No agent can improve the team objective by itself, and there is no smaller objective that belongs to it alone. the Coordinate chapter is largely about living with this.

TT is the episode’s last step, so the sum is finite and γ=1\gamma = 1 would be legitimate here; a continuing task would be written with an infinite limit instead.

Here is the trap, and it is worth stating bluntly because it catches almost everyone.

Every agent receives the same number. That number tells the team how it did. It does not tell any individual agent what it did.

More than three, and the count is not really the point, the kinds are. All three may have chosen badly. Two may have chosen well and one ruined it. Every choice may have been reasonable in isolation and the combination poor anyway. The single number does not distinguish these, because it was never a sum of per-agent contributions in the first place. Nothing was added up, so nothing can be taken apart.

Once you accept that RR is a choice, you have to make it. A reward function is where you write down what “good” means, and most of the difficulty in applying this material is there rather than in the algorithms.

Figure 3
rt=10 1[served]⏟task success+Δpt⏟useful progress−0.5 ct⏟conflicts−0.05⏟time cost\rew = \ubrace{reward}{10\,\mathbb{1}[\text{served}]}{task success} + \ubrace{reward}{\Delta p_t}{useful progress} - \ubrace{conflict}{0.5\,c_t}{conflicts} - \ubrace{conflict}{0.05}{time cost}
1[served]\mathbb{1}[\text{served}]
1 when the order is served on this step, otherwise 0
Δpt\Delta p_t
new task progress, such as completing a required preparation step
ctc_t
the number of collisions or duplicated actions on this step
A shaped kitchen reward that balances completion, progress, conflicts, and time.

The serving bonus encodes the actual task, while the smaller progress terms make useful intermediate behaviour easier to learn. Conflict and time costs discourage waste. Changing any coefficient changes the behaviour that learning favours, so each term needs a reason tied to the objective rather than a value chosen only because it improves a training curve.

Knowledge check

Two kitchen agents act, and the shared reward comes back low. What has each agent learned about its own action?

Select one answer.

  • A cooperative problem gives every agent the same reward rt=R(st,at)\rew = R(\st, \jointact). Their objectives are aligned; their actions are not automatically coordinated.
  • The reward is a function of the joint action, so there is nothing to evaluate for one agent alone.
  • The team maximises one objective J(π)J(\boldsymbol{\pi}), the expected discounted return, and it depends on every agent’s policy at once.
  • A shared reward is not shared credit. It says how the team did, not what any individual did, and an agent that helped can receive the same number as one that did not.