7Partial Observability
Partial observability occurs when different environment states can produce the same local observation, leaving an agent unable to select the correct action from its current input alone.
In this section you will
- Recognize when two states produce the same observation
- Use an observation history to break that ambiguity
- Describe what memory and a belief state add
- Name the uncertainty that is specific to having other agents
The wall is the point. Everything on the other side is still happening, still changing the reward, and still invisible.
Ambiguous Observations
Section titled “Ambiguous Observations”That is the whole problem, and it has a clean formal statement.
- two genuinely different states of the environment
- agent i’s observation function, which discards the part that distinguishes them
Whenever two states collapse to the same observation, an agent that decides from that observation alone must treat them the same way. It has no choice: a function returns one value for one input. If the right action differs between those states, the policy is guaranteed to be wrong in at least one of them, and no amount of training fixes it. The information simply is not there.
Observation Histories
Section titled “Observation Histories”The most useful extra information an agent already has is its own past.
The empty counter is ambiguous now. But an agent that remembers seeing agent 2 walk toward the counter three steps ago has good reason to prefer the first explanation. Nothing about the current observation changed; the interpretation did.
- agent i’s history: everything it has personally seen and done this episode
- what it observed at step k
- what it did at step k, its own actions are part of what it knows
What I saw earlier can help me interpret what I see now.
A policy that reads instead of is called history-dependent; one that reads only the current observation is reactive. Reactive policies are simpler, smaller, and sometimes entirely sufficient. When observations are ambiguous in a way that the past resolves, they are not.
Note what the history is not. It is one agent’s own record. It does not contain what partners saw, and it does not contain the state.
Memory and Belief State
Section titled “Memory and Belief State”Storing a growing list and feeding all of it to a policy does not scale, the history gets longer every step. In practice agents keep a compressed summary instead, and there are three common ways to do it.
Stack the last few observations. Concatenate the most recent observations and treat that as the input. Crude, cheap, and often enough: steps of memory for the cost of a bigger input.
Carry a recurrent state. The policy holds an internal vector, updates it from each new observation, and conditions on it. The network learns what is worth remembering rather than being told. This is what most deep multi-agent methods do, and for our purposes that is the level of detail needed.
Track a belief. Maintain an explicit probability distribution over which state you are in, and update it with each observation. Principled, and expensive: it needs a model of the environment that you usually do not have.
Uncertainty About Other Agents
Section titled “Uncertainty About Other Agents”Everything so far applies to single-agent partial observability. Now the multi-agent part, which is a genuinely different kind of ignorance.
In the single-agent kitchen, one thing was hidden: the state. In the two-agent kitchen there are two, and the second one is not a state at all.
Agent 1 does not know:
- what agent 2 observed. Its partner applied a different observation function to the same state and got something else.
- what agent 2 is about to do. They act simultaneously, so this step’s choice is unavailable by construction.
- what agent 2’s policy is. Even the rule generating those choices is hidden, and during training it is changing.
That third one is worth pausing on. An agent facing an unknown state at least faces a fixed environment. An agent facing an unknown partner policy faces a target that moves as that partner learns, the non-stationarity from From One Agent to Many, arriving here as an information problem rather than a learning-dynamics one.
There are exactly two families of response, and each is a chapter.
Infer it. Build a model of the partner from what you can see of its behaviour, and act on the prediction. This runs through the Adapt chapter, where the question becomes whether your model survives meeting a partner you never trained with.
Be told. Have the partner send you what you are missing. This is the Communicate chapter, and it is not free: a channel has finite capacity, messages can be lost, and a message costs something to send.
Knowledge check
Correct.
Not quite.
A policy that reads only the current observation cannot be right in both cases, so the policy needs memory or a message.
Exactly. A function gives one output per input. If the input is identical and the correct output differs, the input is insufficient, and the fix is more information, not more training.
The policy needs more training on those states.
Training cannot help. The two situations present the same input, so any policy reading only that input must respond identically. The problem is upstream of learning.
The observation function is buggy and should be fixed.
Sometimes true in practice, but not implied. Ambiguity is usually a legitimate feature of the problem: a real agent cannot see through a wall. The framework expects observations to be lossy.
A stochastic policy solves it, since it can pick differently on the two visits.
A stochastic policy will pick differently, but not in a way that tracks which situation it is actually in. It is randomising blindly, so it gets the right action at chance rate. That is sometimes better than being reliably wrong, and it is not knowing.
Explanation
“Same input, different right answer” is the signature of partial observability, and it is worth learning to spot.
Partial Observability Summary
Section titled “Partial Observability Summary”- Partial observability means two different states can produce the same observation: while .
- When that happens, a policy reading only the current observation must act identically in both. No amount of training resolves it, the fix adds information.
- An agent’s own history is the cheapest extra information it has. Policies that use it are history-dependent; those that do not are reactive.
- In practice history is compressed: frame stacking, a recurrent state, or an explicit belief over states.
- Teams add a second unknown: an agent does not know what its partners observed, will do, or have as a policy, and that policy changes while everyone learns.
- Two responses: model the partner (the Adapt chapter) or have it tell you (the Communicate chapter).