6Cooperative Rewards
Cooperative agents pursue one shared objective, represented here by a common reward.
In this section you will
- Contrast per-agent payments with one shared team reward
- Write the reward as a function of the state and the joint action
- State the team objective as an expected discounted return
- Separate aligned objectives from shared credit
- Read a shaped cooperative reward term by term
Individual and Shared Rewards
Section titled “Individual and Shared Rewards”There are really only two options, and choosing between them decides what kind of problem you have.
Pay them separately. Agent 1 earns per vegetable chopped. Agent 2 earns per minute the stove is at temperature. Both can now be measured alone, which is convenient, and immediately wrong. A chef who chops a mountain of onions nobody needs is doing beautifully by that measure. Neither number mentions the order going out.
Pay them once, together. One payment arrives when the dish is served, and both of them receive it. Nobody can earn anything by looking productive. The only way to be paid is for the order to be complete.
The second is the cooperative setting.
The cooperative reward, drawn. One ticket, both agents, paid only when the order is served. Neither agent has a score of its own.
The Shared Reward Function
Section titled “The Shared Reward Function”- the team reward at step t; every agent receives this same value
- the reward function, defined on the state and the joint action together
- the state: the situation the environment is in at step t
- the joint action, which is why no per-agent reward exists
Two features of that line carry the whole section.
It is a function of , the joint action, not of any single agent’s choice. There is no anywhere in the definition, so there is nothing to evaluate for one agent on its own. Ask “what did chopping earn?” and the question has no answer, because the reward function was never given one agent’s action as an input.
And every agent receives the identical value of . In the game-theory vocabulary that makes it a common-reward game, and it is a modelling choice rather than something discovered about the world. You are declaring that this team wins and loses together.
Give every agent the same reward and their objectives are aligned: improving the team return benefits them all. This does not mean their actions will automatically be coordinated. Both kitchen agents want the order served, and they can still both reach for the same pan, or both wait for the other to start. Aligned objectives, interfering actions, the rest of this resource is mostly about that gap.
The Team Objective
Section titled “The Team Objective”One serving is not the goal; a good service is. The team is not chasing the next reward but the discounted sum over the whole episode, and because both the environment and the policies can be random, what gets optimised is its expectation.
- the team objective, a function of the whole joint policy
- the joint policy: the tuple of every agent’s policy
- average over episodes run under that joint policy
- the boxed part: the discounted return of one episode, summed to the episode’s last step T
Read the left-hand side carefully, because it is doing something unusual. There is exactly one objective, and its argument is the whole joint policy .
An agent that improves its own policy while its partners hold still has changed . An agent that improves its own policy while its partners also change may have made worse. No agent can improve the team objective by itself, and there is no smaller objective that belongs to it alone. the Coordinate chapter is largely about living with this.
is the episode’s last step, so the sum is finite and would be legitimate here; a continuing task would be written with an infinite limit instead.
Shared Reward and Credit Assignment
Section titled “Shared Reward and Credit Assignment”Here is the trap, and it is worth stating bluntly because it catches almost everyone.
Every agent receives the same number. That number tells the team how it did. It does not tell any individual agent what it did.
More than three, and the count is not really the point, the kinds are. All three may have chosen badly. Two may have chosen well and one ruined it. Every choice may have been reasonable in isolation and the combination poor anyway. The single number does not distinguish these, because it was never a sum of per-agent contributions in the first place. Nothing was added up, so nothing can be taken apart.
Cooperative Reward Design
Section titled “Cooperative Reward Design”Once you accept that is a choice, you have to make it. A reward function is where you write down what “good” means, and most of the difficulty in applying this material is there rather than in the algorithms.
- 1 when the order is served on this step, otherwise 0
- new task progress, such as completing a required preparation step
- the number of collisions or duplicated actions on this step
The serving bonus encodes the actual task, while the smaller progress terms make useful intermediate behaviour easier to learn. Conflict and time costs discourage waste. Changing any coefficient changes the behaviour that learning favours, so each term needs a reason tied to the objective rather than a value chosen only because it improves a training curve.
Cooperative Rewards Summary
Section titled “Cooperative Rewards Summary”- A cooperative problem gives every agent the same reward . Their objectives are aligned; their actions are not automatically coordinated.
- The reward is a function of the joint action, so there is nothing to evaluate for one agent alone.
- The team maximises one objective , the expected discounted return, and it depends on every agent’s policy at once.
- A shared reward is not shared credit. It says how the team did, not what any individual did, and an agent that helped can receive the same number as one that did not.