Skip to content
MARL in Cooperative Environments
Edit this page

1.4Centralized and Decentralized Learning

5 min read

Centralized and decentralized learning differ in where decisions are made and which information supports them.

In this section you will

  • Describe centralized training with centralized execution
  • Describe decentralized training with decentralized execution
  • Treat training and execution as two separate axes
  • Place a method in the resulting design space

Suppose one learner controls everything. It sees the global state, every agent’s observation, every action taken, and the shared reward. And rather than each agent choosing for itself, this single controller selects the whole joint action.

Figure 1
at=πc(st)\tone{action}{\jointact} = \tone{policy}{\pi_c}\bigl(\tone{observe}{\st}\bigr)
πc\pi_c
a single central policy, the only decision-maker in the system
at\jointact
the entire joint action, chosen at once by that one policy
st\st
the global state, which the controller is assumed to see
One decider, everything visible, choosing for the whole team.

This dissolves the coordination problem. There is no non-stationarity, because nothing else is learning. There is no credit assignment problem, because nothing needs crediting. There is one policy and one reward. And nothing has to be inferred about a partner, because there are no partners.

So why is the rest of this chapter necessary?

The joint action space. With n\nag agents and mm actions each, the controller chooses from

Figure 2
∣A∣=mn\bigl|\mathcal{A}\bigr| = m^{\nag}
∣A∣|\mathcal{A}|
the number of possible joint actions
mm
the number of actions available to each agent
n\nag
the number of agents
A centralized controller chooses from a joint-action space that grows exponentially.

Ten agents with five actions each is nearly ten million options at every single step. The textbook identifies this exponential growth in the joint action space as a central limitation of centralized learning, and it is not a limitation you engineer around. It is the size of the thing being learned.

Physical distribution. Even where the space is small enough, centralized control may simply be unavailable. A central controller needs every observation, in time, every step, and it needs to broadcast every action back. Robots on a warehouse floor, vehicles from different manufacturers, and agents owned by different organizations may not support that assumption before the algorithm is even considered.

The opposite corner. Each agent learns from its own experience and acts on its own observation, with no shared information at any point. This is the fully decentralized form of independent learning.

Figure 3
ati∼πi(ati∣oti)for each i∈N\tone{action}{\act{\ag}} \sim \tone{policy}{\pol{\ag}}\bigl(\act{\ag} \given \tone{observe}{\obs{\ag}}\bigr) \quad \text{for each } \ag \in \mathcal{N}
πi\pol{\ag}
one policy per agent, trained on that agent’s own experience only
oti\obs{\ag}
its own observation, the only input available
Nothing shared, at training or at execution. Scalable, deployable, and blind to partners.

The practical advantages of independent learning still apply: it scales, it deploys, it needs no infrastructure. So does every cost: non-stationarity, and no use of information about the other agents even when that information exists.

Laying the modes out side by side makes the opening visible.

ModeInformation while learningInformation while acting
Centralized training and executioneverythingeverything
Decentralized training and executionlocal onlylocal only
Centralized training, decentralized executioneverythinglocal only

The third row is the one that matters, and notice that it is not a compromise between the first two. It takes the advantage of the first, full information where the learning happens, and the constraint of the second, where the constraint is unavoidable anyway.

That question is the next section.

Knowledge check

A team of eight agents, five actions each, must run on physically separate robots that cannot communicate. Which approach is ruled out, and why?

Select one answer.

Centralized and Decentralized Learning Summary

Section titled “Centralized and Decentralized Learning Summary”
  • Two independent questions: what information is available during training, and what is available during execution.
  • Centralized training and execution removes coordination, credit assignment and non-stationarity in one stroke, and is limited by an exponential joint action space mnm^\nag plus the practical impossibility of central control in distributed systems.
  • Decentralized training and execution, independent learning, scales and deploys, but gives up all information about partners.
  • Local information is a hard constraint at execution and usually no constraint at all during training. That asymmetry is the opening the rest of the chapter exploits.