1.5Centralized Training with Decentralized Execution
Centralized training with decentralized execution, or CTDE, permits global information while policies are learned but restricts each deployed agent to information it can obtain locally.
In this section you will
- Separate what is available during training from what is available at execution
- List the centralized signals that may support learning
- Explain why CTDE is a paradigm rather than an algorithm
- State what CTDE does not guarantee
Training and Execution Phases
Section titled “Training and Execution Phases”During training, centralized information is available and is used to shape the learning signal. During execution, each policy is conditioned only on information locally available to its own agent.
The important word in that sentence is decisions. Nothing is centralized at execution time, not the choice, not the information, not any component. What was centralized is already spent: it went into the gradients that produced these policies, and then it went away.
Centralized Training Information
Section titled “Centralized Training Information”Give the extra information a name so it can be used without re-listing it every time.
- centralized information at step t: available while learning, never at execution
Depending on the setup, might contain
- the global state , if the simulator can provide it,
- other agents’ observations, for ,
- other agents’ actions, or the whole joint action ,
- anything else the trainer happens to have: episode metadata, partner identities, privileged simulator variables.
This is a genuinely free lunch in most workflows, and that is worth being explicit about. Training usually happens in a simulator you wrote, on one machine, with every agent’s tensors in the same process. Declining to look at them does not make your method more principled; it just discards information you already have.
Limits of CTDE
Section titled “Limits of CTDE”This is the correction most worth making early, because CTDE is routinely spoken about as though it were an algorithm.
It is not. CTDE is a training and execution paradigm, a statement about what information is permitted at which moment. It does not tell you:
- whether learning is value-based or policy-based,
- whether you use a centralized critic,
- whether you use value decomposition,
- which algorithm you run at all.
Those are separate design choices made inside CTDE, and the next four sections are exactly those choices. A centralized critic and a value decomposition are both ways of exploiting ; they are siblings under this paradigm rather than alternatives to it.
| Question | Answered by |
|---|---|
| What may training see? What may execution see? | the paradigm, CTDE |
| How is the centralized information actually used? | the method, a centralized critic, a value decomposition, … |
| Which update rule computes it? | the algorithm, PPO, Q-learning, … |
Knowledge check
Correct.
Not quite.
It is not cheating: the extra information shaped the learning signal but is not required to act, so the deployed system still satisfies the decentralization constraint.
Exactly right. The constraint that matters is on execution, and it is satisfied. The global state is a training-time resource, like a simulator or a reset function, things no deployed system has either.
It is cheating, since a real system would not have the global state.
A real system does not have a simulator, a reset button or a million episodes either, and nobody calls those cheating. What matters is whether the *deployed* policy needs information it cannot get, and here it does not.
It is fine, because the global state is only an approximation of the agents’ observations.
It is the other way around: the observations are lossy views of the state. And that is not the reason it is acceptable, the reason is that the state is used only during training.
It depends on whether the agents can communicate at deployment.
That is the CTDE-versus-communication distinction, and it is not what decides this case. Because nothing needs to be transmitted at execution time, communication is irrelevant here, the policies already run on local information alone.
Explanation
Being able to defend this cleanly matters: the objection comes up constantly, and the answer is always about which moment the information is used in.
CTDE Summary
Section titled “CTDE Summary”- CTDE allows centralized information while learning and requires local information while acting.
- names that centralized information: global state, other agents’ observations, the joint action, or anything else training has access to.
- It is a paradigm, not an algorithm. It does not specify value-based versus policy-based learning, nor whether you use a centralized critic or a value decomposition. Those are choices made within it.
- It is not communication. is spent during training; a message is used while acting, and must fit a channel.
- The surviving constraint: deployed policies are functions of local information only.