5Joint Actions and Policies
A joint action combines the individual choices that reach the environment at one time step, while a joint policy combines the local decision rules that produce those choices.
In this section you will
- Write an individual policy that reads only its own observation
- Assemble individual policies into a joint policy
- Compute how a joint-action space grows with the number of agents
- Identify team behaviour that separate local policies cannot represent
Individual Policies
Section titled “Individual Policies”Each agent carries its own policy, and that policy sees only that agent’s observation.
- agent i’s policy: its own decision rule, learned separately
- the only input it gets: this agent’s own observation
- the action it produces
That is the entire decision-making apparatus of one agent. Notice how little it has: no partner’s observation, no partner’s intended action, no view of the state. The chopping agent decides to chop knowing only what is on its own counter.
Joint Actions and Joint Policies
Section titled “Joint Actions and Joint Policies”An individual policy can only be judged by the company it keeps. “Fetch the tomato” is an excellent rule when your partner is heating the pan and a waste of a step when your partner is also fetching the tomato, and the policy itself is identical in both cases.
So the interesting object is not or but what they produce between them.
- the joint policy: the team’s behaviour, in bold
- one agent’s local policy, unchanged and still local
This is the sentence to hold onto: team behaviour is not designed; it is produced. Nobody writes down “agent 1 fetches while agent 2 heats”. Two policies are learned separately, they run separately, and the division of labour is whatever falls out of them.
Build a joint action
Worth doing once so the notation stops being abstract. Note what the function does not do: it makes no decisions. It only records what the agents separately chose.
def joint_action(actions):
"""Combine a list of per-agent actions into one joint action.
actions[i] is the action chosen by agent i + 1.
Return a tuple, so the result is ordered and cannot be modified.
"""
# TODO
return None
print(joint_action([0, 2, 1]))
Output
Hints
- The result should keep the order it was given, but not be modifiable afterwards.
- Python names the conversion function after the type you want. Apply it to `actions` and return that.
One solution
def joint_action(actions):
return tuple(actions)
Other correct answers exist. The checks test behaviour, not wording.
Limits of Separate Decision Rules
Section titled “Limits of Separate Decision Rules”Here is the consequence that makes this more than bookkeeping.
Take reactive local policies, each sampling privately, with nothing correlating the draws. Then the probability of any particular joint action is just the product of the individual probabilities.
- all the observations at step t, one per agent
- agent i's own policy, a distribution over its own actions
- multiply across agents; this needs more than decentralization, namely that each agent samples privately with nothing correlating the draws
The actions are conditionally independent given the observations. Stop conditioning on them and the actions need not be independent at all, because the observations come from one shared state and can carry common information.
Knowledge check
Correct.
Not quite.
Three separate choices among 9, which together land on one of 729 outcomes.
Yes. Nobody chooses among 729. Each agent chooses among its own 9, on its own information, and the joint action is what those three choices amount to when combined.
One choice among 729, made by the joint policy.
That describes a centralized controller, which is a legitimate but different setup. Here the joint policy is a tuple of local policies, and nothing evaluates a mapping from the joint observation to a joint action. A centralized controller removes the need to coordinate separate selections, though it still has to work out which of the 729 is good.
Three separate choices among 729, one per agent.
Each agent’s own action set has 9 elements. 729 is the size of the combined space, which no individual agent selects from.
Three choices among 9 that must be statistically independent of each other.
Close, but the independence is conditional. Given their observations, and assuming each samples privately with nothing correlating the draws, the agents draw independently. Stop conditioning on the observations and the actions need not be independent, since those observations come from one shared state.
Explanation
Holding this distinction makes centralized training with decentralized execution read as an obvious idea rather than a trick: train with everything available, deploy with only what each agent can actually see.
Joint Actions and Policies Summary
Section titled “Joint Actions and Policies Summary”- An individual policy is local: its only input is that agent’s own observation.
- A policy can only be judged in combination. The same rule is good or wasteful depending on what a partner does at the same step.
- The joint policy is the team’s behaviour. It is produced by separate policies, not designed as one.
- Under decentralized execution nothing evaluates a joint mapping. With private independent sampling, the joint distribution is a product.
- That product form is a genuine restriction: randomising which agent takes a role, while guaranteeing exactly one does, needs a varying signal both can read before choosing.