4States, Observations, and Actions
A state records everything relevant that is true in the environment, whereas an observation contains only the information available to one agent.
In this section you will
- Define the environment state
- Write an observation function for one agent
- Identify what a local view discards
- Explain why a decentralized policy conditions on an observation and not on the state
Environment State
Section titled “Environment State”is everything: both agents’ positions, everything they hold, the state of every ingredient, how cooked the pot is, what the order needs.
The environment’s rules are written on . The transition function uses it, the reward function uses it, and nothing in either of them cares whether anybody can see it. The kitchen is in a definite condition regardless of who is looking.
Local Observations
Section titled “Local Observations”Each agent receives its own, generally smaller, view:
- agent i’s observation at step t: what it actually receives
- agent i’s observation function; each agent has its own
- the true state, which the agent never receives directly
For agent 1 in the kitchen, might keep
- the counter directly in front of it,
- what it is currently holding,
- a partner, if that partner is nearby,
- whatever the order ticket displays,
and discards everything else. Note that is a function of the state, so the observation is always genuinely derived from what is true, the agent is not being misled. It is being given less.
State and Observation Diagram
Section titled “State and Observation Diagram”This is the diagram to remember.
The lower panel is not a different picture of the kitchen. It is the same picture, at the same scale, with the same things in the same places, just less of it. That is what an observation function does.
The state is what is true. An observation is what an agent gets to see.
Two consequences follow immediately, and they are the reason this section exists.
Each agent has its own , so partners know different things. In one and the same state, agent 1 may know the tomato is on the counter while agent 2 knows the pot is nearly boiling, and neither knows the other’s fact. There is no shared view to reason from, and there is no agent whose observation is the union of everybody’s. Information in a multi-agent system is distributed, not merely incomplete.
What is unobserved still matters. The hidden half of that diagram is not inert. The stove is still heating, agent 2 is still moving, and the reward at the end of the step will be computed from all of it. An agent is accountable for consequences it cannot see.
Observation-Conditioned Actions
Section titled “Observation-Conditioned Actions”The state-observation distinction determines the information on which an agent can base its action.
- the action agent i takes
- agent i’s own policy
- the only thing the policy is given: this agent’s own observation
Compare this with the single-agent policy from section 1.1. Two substitutions have happened, and both are restrictions:
. The policy is a function of the observation, not the state. It cannot depend on facts the agent does not have.
. There is one policy per agent, not one policy for the system. Nothing evaluates a map from everybody’s observations to everybody’s actions.
Together these are what decentralized execution means, and it is the condition under which every method in this resource has to work. The team’s behaviour has to be produced by separate agents reading separate observations, even if, during training, we allow ourselves to use more than that. Centralized Training with Decentralized Execution is about exactly that loophole.
Knowledge check
Correct.
Not quite.
They are in the same state, but may receive different observations.
Correct. There is exactly one state per step, and each agent applies its own observation function to it. Same truth, different views, and that asymmetry is what makes communication worth studying.
They receive the same observation, since the kitchen is the same.
The kitchen is the same; the views are not. Each agent has its own O_i, and typically they keep different parts. If both agents always received the same observation, most of the Communicate chapter would be unnecessary.
Between them, their observations cover the whole state.
Nothing guarantees that. Both agents can be facing the counter while the pot boils over behind them. Even when their observations do jointly cover the state, no single agent holds that union.
Each can infer the other’s observation from its own.
Only with extra assumptions. An agent that knows the other’s observation function and the state could compute it, but it does not know the state. That is the premise. Estimating what a partner knows is a real technique, covered in Partial Observability, and it is inference rather than access.
Explanation
One state, several views. Holding that shape in mind prevents most of the mistakes in this subject.
States, Observations, and Actions Summary
Section titled “States, Observations, and Actions Summary”- is what is true. is what agent receives. The environment’s rules are written on the first; every agent’s decision is made from the second.
- An observation is a crop of the truth, not a distortion of it.
- Every agent has its own observation function, so partners in the same state know different things and no agent holds the union.
- What an agent cannot see still affects the reward it receives.
- Policies are therefore local: . This is decentralized execution, and it is the constraint every method here must satisfy.