Skip to content
MARL in Cooperative Environments
Edit this page

3The Multi-Agent Environment

7 min read

A multi-agent environment receives decisions from several agents and advances once in response to their combined action.

In this section you will

  • Define the agent set and the environment state
  • Distinguish an individual action from a joint action
  • Trace one state transition driven by a joint action
  • Describe an episode and where it starts

Start by naming them. The set of agents is

Figure 1
N={1,2,…,n}\tone{policy}{\mathcal{N}} = \{1, 2, \dots, \nag\}
N\mathcal{N}
the set of agents, each with its own decision to make
n\nag
how many there are; fixed for the whole episode
A list of deciders. In the kitchen, N = {1, 2}.

Every agent in N\mathcal{N} acts at the same time. At each step all of them choose, then the environment moves once. No agent gets to see what another chose for the current step before choosing itself. There is no turn-taking here, and nothing has been combined yet at the moment of deciding.

The state st∈S\st \in \mathcal{S} is the complete configuration of the environment at step tt. For the kitchen it would include

  • where each agent is standing,
  • what each agent is holding,
  • where the ingredients are, and which are prepared,
  • what is on the stove and how cooked it is,
  • what the current order asks for and how far it has got,
  • how much time is left.

That list is the whole situation, and it is what the environment’s rules are defined on.

Each agent has its own set of things it can do, and picks one:

Figure 2
ati∈Ai\tone{action}{\act{\ag}} \in \mathcal{A}_{\ag}
ati\act{i}
the action agent i takes at step t
Ai\mathcal{A}_{i}
agent i’s own action set; different agents may have different sets
One choice per agent, from that agent's own menu.

In the kitchen the menu is small and physical: move, pick up, put down, prepare, interact with whatever is in front of you. Nothing in that list mentions the other agent. An agent cannot choose “make my partner fetch the tomato”; it can only act, and let the consequences reach the partner through the kitchen.

The environment does not receive two separate actions. It receives one object containing both.

Figure 3
at=(at1, at2, …, atn)\tone{action}{\jointact} = \bigl(\tone{policy}{\act{1}},\ \tone{policy}{\act{2}},\ \dots,\ \tone{policy}{\act{\nag}}\bigr)
at\jointact
the joint action: the whole ordered tuple, in bold
ati\act{i}
agent i’s action, sitting at position i
Everybody's choice for this step, gathered into one input.

The order is not cosmetic: position ii always means agent ii, so agent 1 chops, agent 2 heats is a different tuple from agent 1 heats, agent 2 chops, even though the same two actions appear in both.

And the set of all such tuples is the product of the individual sets:

Figure 4
A=A1×A2×⋯×An\mathcal{A} = \mathcal{A}_{1} \times \mathcal{A}_{2} \times \cdots \times \mathcal{A}_{\nag}
A\mathcal{A}
the joint-action space
Ai\mathcal{A}_i
the local action space available to agent i
n\nag
the number of agents
The joint-action space is the Cartesian product of the local action spaces.

That word product is doing real work. It is worth feeling the size before reading it.

54=6255^4 = 625. Not 5×4=205 \times 4 = 20, which is the count of individual choices available across the team, and not 5+5+5+55 + 5 + 5 + 5. Each agent’s five options multiply against every other agent’s, so the joint action space grows exponentially in the number of agents.

AgentsActions eachJoint actions
2525
45625
85390,625
496,561
8943,046,721

Once the joint action arrives, the environment moves.

Figure 5
st+1∼P(st+1∣st, at)\tone{observe}{s_{t+1}} \sim P\bigl(s_{t+1} \given \tone{observe}{\st},\ \tone{action}{\jointact}\bigr)
st+1s_{t+1}
the next state
PP
the transition function: the environment’s rules
st\st
the state the step began in
at\jointact
the joint action; note P is not defined on any single agent’s action
Where the kitchen ends up depends on what the agents did together.

Read the right-hand side and notice what is missing. There is no P(st+1∣st,at1)P(s_{t+1} \given \st, \act{1}) anywhere in this framework. You cannot ask what agent 1’s action does to the kitchen, because the question is not well-formed: the answer depends on what agent 2 did at the same moment.

Here is the same step drawn out.

One time step. From the state, each agent receives its own observation and selects its own component of the joint action, without seeing the other agents' choices for this step. Those components are combined into a single joint action, which the environment consumes once to produce the next state and, in the shared-reward setting used here, one team reward. each agent seeseach agent choosescombinedenvironmentsto 1a 1o na najoint actionnextstateand one team reward

Each agent decides on its own, on its own information. Those separate decisions are then combined into one input, and the environment transitions once. That single convergence point is the whole geometry of a multi-agent step, and most of this resource is about living with it.

The rest is bookkeeping, and it works as it does in the single-agent case.

An episode starts from an initial state drawn from a distribution μ\mu, a fresh kitchen, a new order. Then the loop above repeats: everybody acts, the state moves, everybody acts again. It ends on a terminal condition, either success (the order is served) or a step limit (service is over). Learning happens across many episodes, not within one.

Knowledge check

You want to know what effect 'place pot on stove' has on the kitchen. Which statement is correct?

Select one answer.

  • N={1,…,n}\mathcal{N} = \{1, \dots, \nag\} agents act simultaneously: no agent conditions on another’s choice for the current step.
  • The state st\st is the whole configuration of the environment. It is not what any agent knows.
  • Each agent picks ati\act{\ag} from its own set; those choices combine into the joint action at\jointact.
  • The joint action space is the product A1×⋯×An\mathcal{A}_1 \times \cdots \times \mathcal{A}_\nag, so it grows exponentially: four agents with five actions each give 625 joint actions.
  • The transition P(st+1∣st,at)P(s_{t+1} \given \st, \jointact) is defined on the joint action. “What does my action do?” has no answer on its own.