Skip to content
MARL in Cooperative Environments
Edit this page

1.5Centralized Training with Decentralized Execution

5 min read

Centralized training with decentralized execution, or CTDE, permits global information while policies are learned but restricts each deployed agent to information it can obtain locally.

In this section you will

  • Separate what is available during training from what is available at execution
  • List the centralized signals that may support learning
  • Explain why CTDE is a paradigm rather than an algorithm
  • State what CTDE does not guarantee
Two panels. In the training panel, centralized information, the global state, all observations and all actions, feeds learning signal, which updates both agents' policies. In the execution panel, each agent receives only its own observation, passes it to its own policy and produces its own action; the space where the centralized information was is empty. TRAININGanything may be usedzglobal state · all observationsall actions · team rewardlearning signalAgent 1 policyAgent 2 policyboth policies updatedEXECUTIONonly what each agent seesno centralized informationo1Agent 1 policya1o2Agent 2 policya2

During training, centralized information is available and is used to shape the learning signal. During execution, each policy is conditioned only on information locally available to its own agent.

The important word in that sentence is decisions. Nothing is centralized at execution time, not the choice, not the information, not any component. What was centralized is already spent: it went into the gradients that produced these policies, and then it went away.

Give the extra information a name so it can be used without re-listing it every time.

Figure 1
zt  =  whatever is available centrally at training time\tone{comm}{z_t} \;=\; \text{whatever is available centrally at training time}
ztz_t
centralized information at step t: available while learning, never at execution
One symbol for everything training is allowed to see and execution is not.

Depending on the setup, ztz_t might contain

  • the global state st\st, if the simulator can provide it,
  • other agents’ observations, otj\obs{j} for j≠ij \neq \ag,
  • other agents’ actions, or the whole joint action at\jointact,
  • anything else the trainer happens to have: episode metadata, partner identities, privileged simulator variables.

This is a genuinely free lunch in most workflows, and that is worth being explicit about. Training usually happens in a simulator you wrote, on one machine, with every agent’s tensors in the same process. Declining to look at them does not make your method more principled; it just discards information you already have.

This is the correction most worth making early, because CTDE is routinely spoken about as though it were an algorithm.

It is not. CTDE is a training and execution paradigm, a statement about what information is permitted at which moment. It does not tell you:

  • whether learning is value-based or policy-based,
  • whether you use a centralized critic,
  • whether you use value decomposition,
  • which algorithm you run at all.

Those are separate design choices made inside CTDE, and the next four sections are exactly those choices. A centralized critic and a value decomposition are both ways of exploiting ztz_t; they are siblings under this paradigm rather than alternatives to it.

QuestionAnswered by
What may training see? What may execution see?the paradigm, CTDE
How is the centralized information actually used?the method, a centralized critic, a value decomposition, …
Which update rule computes it?the algorithm, PPO, Q-learning, …

Knowledge check

A method trains a value function on the global state, then at deployment each agent picks actions using only its own observation. Someone objects that this is 'cheating'. What is the best response?

Select one answer.

  • CTDE allows centralized information while learning and requires local information while acting.
  • ztz_t names that centralized information: global state, other agents’ observations, the joint action, or anything else training has access to.
  • It is a paradigm, not an algorithm. It does not specify value-based versus policy-based learning, nor whether you use a centralized critic or a value decomposition. Those are choices made within it.
  • It is not communication. ztz_t is spent during training; a message is used while acting, and must fit a channel.
  • The surviving constraint: deployed policies are functions of local information only.