Skip to content
MARL in Cooperative Environments
Edit this page

4States, Observations, and Actions

5 min read

A state records everything relevant that is true in the environment, whereas an observation contains only the information available to one agent.

In this section you will

  • Define the environment state
  • Write an observation function for one agent
  • Identify what a local view discards
  • Explain why a decentralized policy conditions on an observation and not on the state

st\st is everything: both agents’ positions, everything they hold, the state of every ingredient, how cooked the pot is, what the order needs.

The environment’s rules are written on st\st. The transition function uses it, the reward function uses it, and nothing in either of them cares whether anybody can see it. The kitchen is in a definite condition regardless of who is looking.

Each agent receives its own, generally smaller, view:

Figure 1
oti=Oi(st)\tone{observe}{\obs{\ag}} = O_{\ag}\bigl(\tone{observe}{\st}\bigr)
oti\obs{i}
agent i’s observation at step t: what it actually receives
OiO_{i}
agent i’s observation function; each agent has its own
st\st
the true state, which the agent never receives directly
An observation is a filter applied to the truth. The agent sees the output, never the input.

For agent 1 in the kitchen, O1O_1 might keep

  • the counter directly in front of it,
  • what it is currently holding,
  • a partner, if that partner is nearby,
  • whatever the order ticket displays,

and discards everything else. Note that O1O_1 is a function of the state, so the observation is always genuinely derived from what is true, the agent is not being misled. It is being given less.

This is the diagram to remember.

Two panels of the same kitchen. The upper panel is the true state and contains Agent 1, a tomato, a counter, a stove, Agent 2 and a plate. The lower panel is what Agent 1 observes: the same left-hand portion of the same scene, containing only Agent 1, the tomato and the counter. The stove, Agent 2 and the plate are still there in the true state but fall outside the observation. TRUE STATEstAgent 1tomatocounterstoveAgent 2plateO1keeps only what this agent can seeWHAT AGENT 1 OBSERVESotAgent 1tomatocounterstill truenot observed

The lower panel is not a different picture of the kitchen. It is the same picture, at the same scale, with the same things in the same places, just less of it. That is what an observation function does.

The state is what is true. An observation is what an agent gets to see.

Two consequences follow immediately, and they are the reason this section exists.

Each agent has its own OiO_i, so partners know different things. In one and the same state, agent 1 may know the tomato is on the counter while agent 2 knows the pot is nearly boiling, and neither knows the other’s fact. There is no shared view to reason from, and there is no agent whose observation is the union of everybody’s. Information in a multi-agent system is distributed, not merely incomplete.

What is unobserved still matters. The hidden half of that diagram is not inert. The stove is still heating, agent 2 is still moving, and the reward at the end of the step will be computed from all of it. An agent is accountable for consequences it cannot see.

The state-observation distinction determines the information on which an agent can base its action.

Figure 2
ati∼πi(ati∣oti)\tone{action}{\act{\ag}} \sim \tone{policy}{\pol{\ag}}\bigl(\act{\ag} \given \tone{observe}{\obs{\ag}}\bigr)
ati\act{i}
the action agent i takes
πi\pol{i}
agent i’s own policy
oti\obs{i}
the only thing the policy is given: this agent’s own observation
Every agent decides alone, from its own partial view.

Compare this with the single-agent policy π(at∣st)\pi(a_t \given \st) from section 1.1. Two substitutions have happened, and both are restrictions:

st→oti\st \rightarrow \obs{\ag}. The policy is a function of the observation, not the state. It cannot depend on facts the agent does not have.

π→πi\pi \rightarrow \pol{\ag}. There is one policy per agent, not one policy for the system. Nothing evaluates a map from everybody’s observations to everybody’s actions.

Together these are what decentralized execution means, and it is the condition under which every method in this resource has to work. The team’s behaviour has to be produced by separate agents reading separate observations, even if, during training, we allow ourselves to use more than that. Centralized Training with Decentralized Execution is about exactly that loophole.

Knowledge check

Two agents are in the same kitchen at the same step. Which of these is guaranteed?

Select one answer.

  • st\st is what is true. oti=Oi(st)\obs{\ag} = O_{\ag}(\st) is what agent i\ag receives. The environment’s rules are written on the first; every agent’s decision is made from the second.
  • An observation is a crop of the truth, not a distortion of it.
  • Every agent has its own observation function, so partners in the same state know different things and no agent holds the union.
  • What an agent cannot see still affects the reward it receives.
  • Policies are therefore local: ati∼πi(⋅∣oti)\act{\ag} \sim \pol{\ag}(\cdot \given \obs{\ag}). This is decentralized execution, and it is the constraint every method here must satisfy.