1.4Centralized and Decentralized Learning
Centralized and decentralized learning differ in where decisions are made and which information supports them.
In this section you will
- Describe centralized training with centralized execution
- Describe decentralized training with decentralized execution
- Treat training and execution as two separate axes
- Place a method in the resulting design space
Centralized Training and Execution
Section titled “Centralized Training and Execution”Suppose one learner controls everything. It sees the global state, every agent’s observation, every action taken, and the shared reward. And rather than each agent choosing for itself, this single controller selects the whole joint action.
- a single central policy, the only decision-maker in the system
- the entire joint action, chosen at once by that one policy
- the global state, which the controller is assumed to see
This dissolves the coordination problem. There is no non-stationarity, because nothing else is learning. There is no credit assignment problem, because nothing needs crediting. There is one policy and one reward. And nothing has to be inferred about a partner, because there are no partners.
So why is the rest of this chapter necessary?
The joint action space. With agents and actions each, the controller chooses from
- the number of possible joint actions
- the number of actions available to each agent
- the number of agents
Ten agents with five actions each is nearly ten million options at every single step. The textbook identifies this exponential growth in the joint action space as a central limitation of centralized learning, and it is not a limitation you engineer around. It is the size of the thing being learned.
Physical distribution. Even where the space is small enough, centralized control may simply be unavailable. A central controller needs every observation, in time, every step, and it needs to broadcast every action back. Robots on a warehouse floor, vehicles from different manufacturers, and agents owned by different organizations may not support that assumption before the algorithm is even considered.
Decentralized Training and Execution
Section titled “Decentralized Training and Execution”The opposite corner. Each agent learns from its own experience and acts on its own observation, with no shared information at any point. This is the fully decentralized form of independent learning.
- one policy per agent, trained on that agent’s own experience only
- its own observation, the only input available
The practical advantages of independent learning still apply: it scales, it deploys, it needs no infrastructure. So does every cost: non-stationarity, and no use of information about the other agents even when that information exists.
Training and Execution as Separate Axes
Section titled “Training and Execution as Separate Axes”Laying the modes out side by side makes the opening visible.
| Mode | Information while learning | Information while acting |
|---|---|---|
| Centralized training and execution | everything | everything |
| Decentralized training and execution | local only | local only |
| Centralized training, decentralized execution | everything | local only |
The third row is the one that matters, and notice that it is not a compromise between the first two. It takes the advantage of the first, full information where the learning happens, and the constraint of the second, where the constraint is unavoidable anyway.
That question is the next section.
Knowledge check
Correct.
Not quite.
Centralized execution, twice over: the joint action space has 390,625 entries, and the robots could not reach a central controller anyway.
Both reasons are independently fatal, and it is worth seeing that they are separate. Even with a tractable action space the communication requirement would sink it, and even with perfect communication the exponential space would.
Centralized training, the agents cannot share information.
This is the confusion the section exists to prevent. The robots cannot share information at execution time. Training almost certainly happens in a simulator beforehand, where every observation is available. The deployment constraint says nothing about the training constraint.
Decentralized execution, eight agents cannot coordinate without communication.
They can, and often must. Decentralized execution is the requirement here, not the thing ruled out. Coordinating without communication is harder, and the Communicate chapter is about what changes when a channel is available.
Nothing is ruled out; all three modes remain available.
Centralized execution needs every observation delivered to one controller each step and every action sent back. With robots that cannot communicate, that is not an option.
Explanation
Keeping “what is available while learning” and “what is available while acting” as two separate questions is the single most useful habit in this chapter.
Centralized and Decentralized Learning Summary
Section titled “Centralized and Decentralized Learning Summary”- Two independent questions: what information is available during training, and what is available during execution.
- Centralized training and execution removes coordination, credit assignment and non-stationarity in one stroke, and is limited by an exponential joint action space plus the practical impossibility of central control in distributed systems.
- Decentralized training and execution, independent learning, scales and deploys, but gives up all information about partners.
- Local information is a hard constraint at execution and usually no constraint at all during training. That asymmetry is the opening the rest of the chapter exploits.