Skip to content
MARL in Cooperative Environments
Edit this page

5Joint Actions and Policies

6 min read

A joint action combines the individual choices that reach the environment at one time step, while a joint policy combines the local decision rules that produce those choices.

In this section you will

  • Write an individual policy that reads only its own observation
  • Assemble individual policies into a joint policy
  • Compute how a joint-action space grows with the number of agents
  • Identify team behaviour that separate local policies cannot represent

Each agent carries its own policy, and that policy sees only that agent’s observation.

Figure 1
ati∼πi(ati∣oti)\tone{action}{\act{\ag}} \sim \tone{policy}{\pol{\ag}}\bigl(\act{\ag} \given \tone{observe}{\obs{\ag}}\bigr)
πi\pol{i}
agent i’s policy: its own decision rule, learned separately
oti\obs{i}
the only input it gets: this agent’s own observation
ati\act{i}
the action it produces
One question, asked privately by each agent: given what I know, what should I do?

That is the entire decision-making apparatus of one agent. Notice how little it has: no partner’s observation, no partner’s intended action, no view of the state. The chopping agent decides to chop knowing only what is on its own counter.

An individual policy can only be judged by the company it keeps. “Fetch the tomato” is an excellent rule when your partner is heating the pan and a waste of a step when your partner is also fetching the tomato, and the policy itself is identical in both cases.

So the interesting object is not π1\pol{1} or π2\pol{2} but what they produce between them.

Figure 2
π=(π1,π2,…,πn)\boldsymbol{\pi} = \bigl(\tone{policy}{\pol{1}}, \tone{policy}{\pol{2}}, \dots, \tone{policy}{\pol{\nag}}\bigr)
π\boldsymbol{\pi}
the joint policy: the team’s behaviour, in bold
πi\pol{i}
one agent’s local policy, unchanged and still local
A team's behaviour is a bundle of separate decision rules, not one big decision rule.

This is the sentence to hold onto: team behaviour is not designed; it is produced. Nobody writes down “agent 1 fetches while agent 2 heats”. Two policies are learned separately, they run separately, and the division of labour is whatever falls out of them.

Exercise

Build a joint action

Worth doing once so the notation stops being abstract. Note what the function does not do: it makes no decisions. It only records what the agents separately chose.

def joint_action(actions):
    """Combine a list of per-agent actions into one joint action.

    actions[i] is the action chosen by agent i + 1.
    Return a tuple, so the result is ordered and cannot be modified.
    """
    # TODO
    return None


print(joint_action([0, 2, 1]))

Here is the consequence that makes this more than bookkeeping.

Take reactive local policies, each sampling privately, with nothing correlating the draws. Then the probability of any particular joint action is just the product of the individual probabilities.

Figure 3
Pr⁡(at∣ot)=∏i=1nπi(ati∣oti)⏟one agent, local input\Pr\bigl(\tone{action}{\jointact} \given \tone{observe}{\mathbf{o}_t}\bigr) = \prod_{i=1}^{\nag} \ubrace{policy}{\pol{i}\bigl(\act{i} \given \obs{i}\bigr)}{one agent, local input}
ot\mathbf{o}_t
all the observations at step t, one per agent
πi\pol{i}
agent i's own policy, a distribution over its own actions
∏i=1n\prod_{i=1}^{\nag}
multiply across agents; this needs more than decentralization, namely that each agent samples privately with nothing correlating the draws
Everyone deciding alone, at the same moment, multiplies out.

The actions are conditionally independent given the observations. Stop conditioning on them and the actions need not be independent at all, because the observations come from one shared state and can carry common information.

Knowledge check

Three agents, 9 actions each. Which best describes what the joint policy is doing at one step?

Select one answer.

  • An individual policy πi(ati∣oti)\pol{\ag}(\act{\ag} \given \obs{\ag}) is local: its only input is that agent’s own observation.
  • A policy can only be judged in combination. The same rule is good or wasteful depending on what a partner does at the same step.
  • The joint policy π=(π1,…,πn)\boldsymbol{\pi} = (\pol{1}, \dots, \pol{\nag}) is the team’s behaviour. It is produced by separate policies, not designed as one.
  • Under decentralized execution nothing evaluates a joint mapping. With private independent sampling, the joint distribution is a product.
  • That product form is a genuine restriction: randomising which agent takes a role, while guaranteeing exactly one does, needs a varying signal both can read before choosing.