Skip to content
MARL in Cooperative Environments
Edit this page

8The Dec-POMDP Framework

6 min read

A decentralized partially observable Markov decision process, or Dec-POMDP, is the standard model for cooperative agents that act from local information and receive a shared reward.

In this section you will

  • List the components of a Dec-POMDP
  • Follow one complete interaction step through the model
  • State the joint-policy objective
  • Map the cooperative kitchen into the framework
Figure 1
M=⟨ N, S, {Ai}, P, R, {Ωi}, {Oi}, μ, γ ⟩\mathcal{M} = \bigl\langle\, \tone{policy}{\mathcal{N}},\ \tone{observe}{\mathcal{S}},\ \tone{action}{\{\mathcal{A}_\ag\}},\ P,\ \tone{reward}{R},\ \tone{observe}{\{\Omega_\ag\}},\ \tone{observe}{\{O_\ag\}},\ \mu,\ \gamma \,\bigr\rangle
N\mathcal{N}
the agents, the deciders
S\mathcal{S}
the states, what can be true
{Ai}\{\mathcal{A}_\ag\}
each agent’s own action set
PP
the transition function, defined on the joint action
RR
the one reward function every agent shares
{Ωi}\{\Omega_\ag\}
each agent’s set of possible observations
{Oi}\{O_\ag\}
each agent’s observation function, mapping a state to what it sees
μ\mu
the distribution episodes start from
γ\gamma
the discount factor
A decentralized partially observable Markov decision process. Everything you have met, in one tuple.

That is a Dec-POMDP, and it is the standard model for a cooperative multi-agent problem. Presentations of it vary in small ways, some fold μ\mu into S\mathcal{S}, some write a single observation function instead of one per agent, so expect minor differences between papers rather than assuming you have misread one.

Read the acronym backwards and it tells you its own history.

MDP. A Markov decision process: states, actions, a transition rule, a reward, a discount. Sutton and Barto’s subject, and Reinforcement Learning.

PO-MDP. Partially observable. Add observation sets and observation functions, and the agent no longer sees the state.

Dec-POMDP. Decentralized. Add more agents, each with its own actions and its own observations, all sharing one reward function, and crucially, no central decider. The team’s behaviour is produced by separate policies.

Here is the whole model running for a single step.

One time step. From the state, each agent receives its own observation and selects its own component of the joint action, without seeing the other agents' choices for this step. Those components are combined into a single joint action, which the environment consumes once to produce the next state and, in the shared-reward setting used here, one team reward. each agent seeseach agent choosescombinedenvironmentsto 1a 1⋮⋮o na najoint actionnextstateand one team reward

Follow it left to right. There is one state. Each agent’s observation function produces its own observation from that state. Each agent’s policy turns its own observation into its own action, with no access to the others’ choices. Those actions combine into one joint action. The environment consumes it once, and returns a next state and a single team reward that everybody receives.

Every difficulty in this resource is visible in that picture. The narrowing at each agent’s observation is partial observability. The convergence at the joint action is interdependence. The single reward coming back out, undivided, is credit assignment.

The goal is stated on the joint policy, because there is nothing smaller to state it on.

Figure 2
π∗=arg⁡max⁡π  Eπ[  ∑t=0Tγt rt  ]\boldsymbol{\pi}^{*} = \arg\max_{\boldsymbol{\pi}} \; \E_{\boldsymbol{\pi}}\Bigl[\; \cbox{reward}{\sum_{t=0}^{T} \gamma^{t}\, \rew} \;\Bigr]
π∗\boldsymbol{\pi}^{*}
the best joint policy: one local policy per agent
arg⁡max⁡π\arg\max_{\boldsymbol{\pi}}
maximise over the whole bundle at once, not one agent at a time
∑γtrt\sum \gamma^{t} \rew
the discounted team return of an episode
One objective, maximised over every agent's policy simultaneously.

Note what is being searched over. Not a single policy but a tuple of them, and every element must be a function of its own agent’s local information only. So the problem the rest of this resource works on is:

Find decentralized policies that produce strong team behaviour despite interdependent actions and incomplete information.

Both halves of that sentence are constraints. Drop “decentralized” and it becomes an ordinary, if large, single-agent problem. Drop “incomplete information” and coordination gets dramatically easier.

The model is abstract; the kitchen has been concrete throughout. Here they are side by side, which is the point of having carried one example the whole way.

PieceIn the kitchen
N\mathcal{N}{1,2}\{1, 2\}, two agents
S\mathcal{S}every configuration of positions, ingredients, stove and order
Ai\mathcal{A}_\agmove, pick up, put down, prepare, interact
PPwhat the kitchen does given both agents’ actions
RRone payment when an order is served, to both agents
Ωi\Omega_\agthe views an agent can have of its own surroundings
OiO_\agthe crop: nearby counter, what it holds, a nearby partner, the ticket
μ\mua fresh kitchen with a new order
γ\gammahow much a served order later is worth now

Nothing in the left column was invented to fit the right. The kitchen was always this object; we just had not written it down.

Knowledge check

Which single change would turn a Dec-POMDP into an ordinary single-agent POMDP?

Select one answer.

  • A cooperative multi-agent problem is a Dec-POMDP: agents, states, per-agent actions, a joint transition rule, one shared reward, per-agent observations and observation functions, a start distribution and a discount.
  • The name decomposes: MDP → partially observable → decentralized.
  • One step: one state, one observation per agent, one action per agent, one joint action, one transition, one shared reward.
  • The objective is over the joint policy, with every component restricted to local information: find decentralized policies that work well together despite interdependence and incomplete information.
  • Different cooperative domains can share the same Dec-POMDP structure.

The standard reference is Multi-Agent Reinforcement Learning: Foundations and Modern Approaches by Stefano V. Albrecht, Filippos Christianos and Lukas Schäfer (MIT Press, 2024), free at marl-book.com. Its Chapter 3 develops the game models this section places us among, and its Chapter 4 covers the solution concepts the Coordinate chapter will need.

For a compact survey of the same ground alongside single-agent RL, Chapter 5 of Reinforcement Learning: A Comprehensive Overview by Kevin P. Murphy (arXiv:2412.05265) covers multi-agent RL in eighteen pages, with a figure of the game hierarchy that is the quickest single picture of where a Dec-POMDP sits.