Skip to content
MARL in Cooperative Environments
Edit this page

7Partial Observability

6 min read

Partial observability occurs when different environment states can produce the same local observation, leaving an agent unable to select the correct action from its current input alone.

In this section you will

  • Recognize when two states produce the same observation
  • Use an observation history to break that ambiguity
  • Describe what memory and a belief state add
  • Name the uncertainty that is specific to having other agents
Two robot agents working in one kitchen. The left agent is at a chopping board, the right agent at a stove with a pot. A partition stands between them, so neither can see what the other is doing. Agent 1Agent 2neither can see past this

The wall is the point. Everything on the other side is still happening, still changing the reward, and still invisible.

That is the whole problem, and it has a clean formal statement.

Figure 1
s(1)≠s(2)butOi(s(1))=Oi(s(2))\tone{observe}{s^{(1)}} \neq \tone{observe}{s^{(2)}} \qquad\text{but}\qquad O_{\ag}\bigl(\tone{observe}{s^{(1)}}\bigr) = O_{\ag}\bigl(\tone{observe}{s^{(2)}}\bigr)
s(1),s(2)s^{(1)}, s^{(2)}
two genuinely different states of the environment
OiO_{i}
agent i’s observation function, which discards the part that distinguishes them
Two different truths that look identical to this agent.

Whenever two states collapse to the same observation, an agent that decides from that observation alone must treat them the same way. It has no choice: a function returns one value for one input. If the right action differs between those states, the policy is guaranteed to be wrong in at least one of them, and no amount of training fixes it. The information simply is not there.

The most useful extra information an agent already has is its own past.

The empty counter is ambiguous now. But an agent that remembers seeing agent 2 walk toward the counter three steps ago has good reason to prefer the first explanation. Nothing about the current observation changed; the interpretation did.

Figure 2
hti=(o0i, a0i, o1i, a1i, …, oti)\tone{observe}{h^{\ag}_t} = \bigl(\tone{observe}{o^{\ag}_0},\ \tone{action}{a^{\ag}_0},\ \tone{observe}{o^{\ag}_1},\ \tone{action}{a^{\ag}_1},\ \dots,\ \tone{observe}{\obs{\ag}}\bigr)
htih^{i}_t
agent i’s history: everything it has personally seen and done this episode
okio^{i}_k
what it observed at step k
akia^{i}_k
what it did at step k, its own actions are part of what it knows
One agent's private record of the episode so far.

What I saw earlier can help me interpret what I see now.

A policy that reads htih^{\ag}_t instead of oti\obs{\ag} is called history-dependent; one that reads only the current observation is reactive. Reactive policies are simpler, smaller, and sometimes entirely sufficient. When observations are ambiguous in a way that the past resolves, they are not.

Note what the history is not. It is one agent’s own record. It does not contain what partners saw, and it does not contain the state.

Storing a growing list and feeding all of it to a policy does not scale, the history gets longer every step. In practice agents keep a compressed summary instead, and there are three common ways to do it.

Stack the last few observations. Concatenate the most recent kk observations and treat that as the input. Crude, cheap, and often enough: kk steps of memory for the cost of a bigger input.

Carry a recurrent state. The policy holds an internal vector, updates it from each new observation, and conditions on it. The network learns what is worth remembering rather than being told. This is what most deep multi-agent methods do, and for our purposes that is the level of detail needed.

Track a belief. Maintain an explicit probability distribution over which state you are in, and update it with each observation. Principled, and expensive: it needs a model of the environment that you usually do not have.

Everything so far applies to single-agent partial observability. Now the multi-agent part, which is a genuinely different kind of ignorance.

In the single-agent kitchen, one thing was hidden: the state. In the two-agent kitchen there are two, and the second one is not a state at all.

Agent 1 does not know:

  • what agent 2 observed. Its partner applied a different observation function to the same state and got something else.
  • what agent 2 is about to do. They act simultaneously, so this step’s choice is unavailable by construction.
  • what agent 2’s policy is. Even the rule generating those choices is hidden, and during training it is changing.

That third one is worth pausing on. An agent facing an unknown state at least faces a fixed environment. An agent facing an unknown partner policy faces a target that moves as that partner learns, the non-stationarity from From One Agent to Many, arriving here as an information problem rather than a learning-dynamics one.

There are exactly two families of response, and each is a chapter.

Infer it. Build a model of the partner from what you can see of its behaviour, and act on the prediction. This runs through the Adapt chapter, where the question becomes whether your model survives meeting a partner you never trained with.

Be told. Have the partner send you what you are missing. This is the Communicate chapter, and it is not free: a channel has finite capacity, messages can be lost, and a message costs something to send.

Knowledge check

An agent's two observations at different steps are byte-for-byte identical, but the best action differs between them. What follows?

Select one answer.

  • Partial observability means two different states can produce the same observation: s(1)≠s(2)s^{(1)} \neq s^{(2)} while Oi(s(1))=Oi(s(2))O_\ag(s^{(1)}) = O_\ag(s^{(2)}).
  • When that happens, a policy reading only the current observation must act identically in both. No amount of training resolves it, the fix adds information.
  • An agent’s own history htih^\ag_t is the cheapest extra information it has. Policies that use it are history-dependent; those that do not are reactive.
  • In practice history is compressed: frame stacking, a recurrent state, or an explicit belief over states.
  • Teams add a second unknown: an agent does not know what its partners observed, will do, or have as a policy, and that policy changes while everyone learns.
  • Two responses: model the partner (the Adapt chapter) or have it tell you (the Communicate chapter).