3The Multi-Agent Environment
A multi-agent environment receives decisions from several agents and advances once in response to their combined action.
In this section you will
- Define the agent set and the environment state
- Distinguish an individual action from a joint action
- Trace one state transition driven by a joint action
- Describe an episode and where it starts
The Agent Set
Section titled “The Agent Set”Start by naming them. The set of agents is
- the set of agents, each with its own decision to make
- how many there are; fixed for the whole episode
Every agent in acts at the same time. At each step all of them choose, then the environment moves once. No agent gets to see what another chose for the current step before choosing itself. There is no turn-taking here, and nothing has been combined yet at the moment of deciding.
Environment State
Section titled “Environment State”The state is the complete configuration of the environment at step . For the kitchen it would include
- where each agent is standing,
- what each agent is holding,
- where the ingredients are, and which are prepared,
- what is on the stove and how cooked it is,
- what the current order asks for and how far it has got,
- how much time is left.
That list is the whole situation, and it is what the environment’s rules are defined on.
Individual Actions
Section titled “Individual Actions”Each agent has its own set of things it can do, and picks one:
- the action agent i takes at step t
- agent i’s own action set; different agents may have different sets
In the kitchen the menu is small and physical: move, pick up, put down, prepare, interact with whatever is in front of you. Nothing in that list mentions the other agent. An agent cannot choose “make my partner fetch the tomato”; it can only act, and let the consequences reach the partner through the kitchen.
Joint Actions
Section titled “Joint Actions”The environment does not receive two separate actions. It receives one object containing both.
- the joint action: the whole ordered tuple, in bold
- agent i’s action, sitting at position i
The order is not cosmetic: position always means agent , so agent 1 chops, agent 2 heats is a different tuple from agent 1 heats, agent 2 chops, even though the same two actions appear in both.
And the set of all such tuples is the product of the individual sets:
- the joint-action space
- the local action space available to agent i
- the number of agents
That word product is doing real work. It is worth feeling the size before reading it.
. Not , which is the count of individual choices available across the team, and not . Each agent’s five options multiply against every other agent’s, so the joint action space grows exponentially in the number of agents.
| Agents | Actions each | Joint actions |
|---|---|---|
| 2 | 5 | 25 |
| 4 | 5 | 625 |
| 8 | 5 | 390,625 |
| 4 | 9 | 6,561 |
| 8 | 9 | 43,046,721 |
State Transitions
Section titled “State Transitions”Once the joint action arrives, the environment moves.
- the next state
- the transition function: the environment’s rules
- the state the step began in
- the joint action; note P is not defined on any single agent’s action
Read the right-hand side and notice what is missing. There is no anywhere in this framework. You cannot ask what agent 1’s action does to the kitchen, because the question is not well-formed: the answer depends on what agent 2 did at the same moment.
Here is the same step drawn out.
Each agent decides on its own, on its own information. Those separate decisions are then combined into one input, and the environment transitions once. That single convergence point is the whole geometry of a multi-agent step, and most of this resource is about living with it.
Episodes and Initial States
Section titled “Episodes and Initial States”The rest is bookkeeping, and it works as it does in the single-agent case.
An episode starts from an initial state drawn from a distribution , a fresh kitchen, a new order. Then the loop above repeats: everybody acts, the state moves, everybody acts again. It ends on a terminal condition, either success (the order is served) or a step limit (service is over). Learning happens across many episodes, not within one.
Knowledge check
Correct.
Not quite.
The question is not well-formed on its own. The transition is defined on the joint action, so the effect depends on what the other agent does at the same step.
Exactly. P takes the joint action as input. You can ask what a full joint action does; you cannot ask what one agent’s component of it does without fixing the rest.
The pot ends up on the stove, since that action affects only the pot.
Usually, but not reliably, and the reliability is the point. If the other agent is carrying the pot away at the same step, or is occupying the stove, the outcome differs. The framework does not let you assume an action’s effect is local.
You can answer it by averaging over the other agent’s actions.
You can, and that gives a useful summary, but only once you have chosen a distribution over what the other agent does, which means you have assumed a partner policy. That is a real technique, not a property of the environment, and it is exactly the assumption that goes stale when partners learn.
It depends on the reward function.
The reward and the transition are separate objects. The reward scores what happened; the transition determines what happens. Both take the joint action, but the reward has no say in where the kitchen ends up.
Explanation
Being able to say “that question needs the joint action” is what separates a single-agent habit from a multi-agent one.
Multi-Agent Environment Summary
Section titled “Multi-Agent Environment Summary”- agents act simultaneously: no agent conditions on another’s choice for the current step.
- The state is the whole configuration of the environment. It is not what any agent knows.
- Each agent picks from its own set; those choices combine into the joint action .
- The joint action space is the product , so it grows exponentially: four agents with five actions each give 625 joint actions.
- The transition is defined on the joint action. “What does my action do?” has no answer on its own.