8The Dec-POMDP Framework
A decentralized partially observable Markov decision process, or Dec-POMDP, is the standard model for cooperative agents that act from local information and receive a shared reward.
In this section you will
- List the components of a Dec-POMDP
- Follow one complete interaction step through the model
- State the joint-policy objective
- Map the cooperative kitchen into the framework
Dec-POMDP Components
Section titled “Dec-POMDP Components”- the agents, the deciders
- the states, what can be true
- each agent’s own action set
- the transition function, defined on the joint action
- the one reward function every agent shares
- each agent’s set of possible observations
- each agent’s observation function, mapping a state to what it sees
- the distribution episodes start from
- the discount factor
That is a Dec-POMDP, and it is the standard model for a cooperative multi-agent problem. Presentations of it vary in small ways, some fold into , some write a single observation function instead of one per agent, so expect minor differences between papers rather than assuming you have misread one.
Read the acronym backwards and it tells you its own history.
MDP. A Markov decision process: states, actions, a transition rule, a reward, a discount. Sutton and Barto’s subject, and Reinforcement Learning.
PO-MDP. Partially observable. Add observation sets and observation functions, and the agent no longer sees the state.
Dec-POMDP. Decentralized. Add more agents, each with its own actions and its own observations, all sharing one reward function, and crucially, no central decider. The team’s behaviour is produced by separate policies.
A Complete Dec-POMDP Step
Section titled “A Complete Dec-POMDP Step”Here is the whole model running for a single step.
Follow it left to right. There is one state. Each agent’s observation function produces its own observation from that state. Each agent’s policy turns its own observation into its own action, with no access to the others’ choices. Those actions combine into one joint action. The environment consumes it once, and returns a next state and a single team reward that everybody receives.
Every difficulty in this resource is visible in that picture. The narrowing at each agent’s observation is partial observability. The convergence at the joint action is interdependence. The single reward coming back out, undivided, is credit assignment.
The Joint-Policy Objective
Section titled “The Joint-Policy Objective”The goal is stated on the joint policy, because there is nothing smaller to state it on.
- the best joint policy: one local policy per agent
- maximise over the whole bundle at once, not one agent at a time
- the discounted team return of an episode
Note what is being searched over. Not a single policy but a tuple of them, and every element must be a function of its own agent’s local information only. So the problem the rest of this resource works on is:
Find decentralized policies that produce strong team behaviour despite interdependent actions and incomplete information.
Both halves of that sentence are constraints. Drop “decentralized” and it becomes an ordinary, if large, single-agent problem. Drop “incomplete information” and coordination gets dramatically easier.
The Kitchen as a Dec-POMDP
Section titled “The Kitchen as a Dec-POMDP”The model is abstract; the kitchen has been concrete throughout. Here they are side by side, which is the point of having carried one example the whole way.
| Piece | In the kitchen |
|---|---|
| , two agents | |
| every configuration of positions, ingredients, stove and order | |
| move, pick up, put down, prepare, interact | |
| what the kitchen does given both agents’ actions | |
| one payment when an order is served, to both agents | |
| the views an agent can have of its own surroundings | |
| the crop: nearby counter, what it holds, a nearby partner, the ticket | |
| a fresh kitchen with a new order | |
| how much a served order later is worth now |
Nothing in the left column was invented to fit the right. The kitchen was always this object; we just had not written it down.
Knowledge check
Correct.
Not quite.
One controller receives every agent’s observation and selects the whole joint action.
Correct. With one decider holding all the observations and choosing the joint action, there is a single policy over a single (large) action space, and the decentralization constraint is gone. This is the central-learning reduction, and it is exactly what *Centralized and Decentralized Learning* examines, along with why the exponential joint action space makes it impractical.
Every agent is given the full state.
That removes the partial observability, giving a multi-agent MDP. It does not remove the multi-agent part: several agents still choose separately and simultaneously, so coordination is still required.
Reduce the number of agents to two.
Two agents is still multi-agent, and this chapter has used two throughout precisely because two is already enough for every difficulty to appear. Agent count changes the size of the joint action space, not the kind of problem.
Give each agent its own reward function.
That goes the other way: it drops the common reward and gives you a partially observable stochastic game, which is a more general and harder setting rather than a single-agent one.
Explanation
Knowing what would make the problem easy is a good way to see what actually makes it hard.
Dec-POMDP Summary
Section titled “Dec-POMDP Summary”- A cooperative multi-agent problem is a Dec-POMDP: agents, states, per-agent actions, a joint transition rule, one shared reward, per-agent observations and observation functions, a start distribution and a discount.
- The name decomposes: MDP → partially observable → decentralized.
- One step: one state, one observation per agent, one action per agent, one joint action, one transition, one shared reward.
- The objective is over the joint policy, with every component restricted to local information: find decentralized policies that work well together despite interdependence and incomplete information.
- Different cooperative domains can share the same Dec-POMDP structure.
Further reading
Section titled “Further reading”The standard reference is Multi-Agent Reinforcement Learning: Foundations and Modern Approaches by Stefano V. Albrecht, Filippos Christianos and Lukas Schäfer (MIT Press, 2024), free at marl-book.com. Its Chapter 3 develops the game models this section places us among, and its Chapter 4 covers the solution concepts the Coordinate chapter will need.
For a compact survey of the same ground alongside single-agent RL, Chapter 5 of Reinforcement Learning: A Comprehensive Overview by Kevin P. Murphy (arXiv:2412.05265) covers multi-agent RL in eighteen pages, with a figure of the game hierarchy that is the quickest single picture of where a Dec-POMDP sits.