Skip to content
MARL in Cooperative Environments
Edit this page

1.6Centralized Critics

5 min read

A centralized critic is one way to use global information under CTDE: the critic evaluates actions during training while each actor remains a local policy for execution.

In this section you will

  • Separate the roles of the actor and the critic
  • Expand the critic’s inputs to the state and the joint action
  • Explain how a richer evaluation changes the policy update
  • State what remains decentralized at execution

Many reinforcement learning methods split the learner in two.

The actor is the policy: the thing that chooses actions. The critic is a learned value or advantage function used to improve that policy. It does not choose anything, it evaluates.

Figure 1
πθ(a∣s)choosesVϕ(s)evaluates\underset{\text{\scriptsize chooses}}{\tone{policy}{\pi_\theta(a \given s)}} \qquad\qquad \underset{\text{\scriptsize evaluates}}{\tone{reward}{V_\phi(s)}}
πθ\pi_\theta
the actor: the policy, with its own parameters θ
VϕV_\phi
the critic: a learned estimate of how good a situation is, parameters φ
Two components with two jobs. Only one of them has to run at deployment.

The critic exists to make the actor’s updates less noisy. Instead of judging an action by the whole messy return that followed it, the actor is nudged according to how that action compared with what the critic expected.

And here is the observation that makes this section work: only the actor is needed at deployment. The critic is training scaffolding. Once the policy is trained, the critic is discarded.

That is the whole idea. Condition the critic on the centralized information ztz_t, and leave the actor exactly as constrained as it has to be.

Figure 2
Vi(hti, zt)orQi(hti, zt, at)\tone{reward}{V_\ag\bigl(\tone{observe}{h^\ag_t},\ \tone{comm}{z_t}\bigr)} \qquad\text{or}\qquad \tone{reward}{Q_\ag\bigl(\tone{observe}{h^\ag_t},\ \tone{comm}{z_t},\ \tone{action}{\jointact}\bigr)}
htih^\ag_t
agent i’s own history, what it has personally seen and done
ztz_t
centralized information, available because this is training
at\jointact
the joint action; an action-value critic is conditioned on what everybody did
Two flavours of centralized critic: one evaluates the situation, one evaluates the situation together with what the whole team did.

Both forms appear in the literature. A centralized state-value critic is conditioned on the agent’s history plus the centralized information. A centralized action-value critic is additionally conditioned on the joint action, which, from Coordination, is exactly the form that can tell combinations apart.

Meanwhile the actor is untouched:

Figure 3
πi(ati∣hti)\tone{policy}{\pol{\ag}\bigl(\tone{action}{\act{\ag}} \given \tone{observe}{h^\ag_t}\bigr)}
πi\pol{\ag}
the actor: still local, still deployable, still the only thing that runs
htih^\ag_t
its own history and nothing else, no z, no partner action
The part that survives deployment never learned to depend on anything it will not have.
Two panels. In the training panel, centralized information, the global state, all observations and all actions, feeds critic, which updates both agents' policies. In the execution panel, each agent receives only its own observation, passes it to its own policy and produces its own action; the space where the centralized information was is empty. TRAININGanything may be usedzglobal state · all observationsall actions · team rewardcriticAgent 1 policyAgent 2 policyboth policies updatedEXECUTIONonly what each agent seesno centralized informationo1Agent 1 policya1o2Agent 2 policya2

The picture is the CTDE picture with one box renamed, and that is the point. A centralized critic is not a different arrangement. It is a specific answer to “what consumes ztz_t?”

The critic’s job is to say how good things were. A blind critic has to do that from one agent’s partial view, which means the same observation gets the same evaluation whether the team’s other half was doing something useful or something disastrous. Its estimates are noisy for a reason that has nothing to do with the agent being evaluated.

Give it ztz_t and it can see what the team actually did. Its estimates get sharper, the actor’s updates get less noisy, and learning gets more stable, all without changing what the actor is allowed to know.

Knowledge check

Why is it acceptable for a critic to use the global state when the actor may not?

Select one answer.

  • An actor chooses actions; a critic evaluates. Only the actor is needed at deployment.
  • A centralized critic is conditioned on the agent’s history plus centralized information ztz_t; an action-value version is additionally conditioned on the joint action at\jointact.
  • The actor stays local: πi(ati∣hti)\pol{\ag}(\act{\ag} \given h^\ag_t).
  • The benefit is a better-informed learning signal, the critic can see what the team did, so its estimates are less confounded by partner behaviour.
  • It does not solve coordination. The actor still decides alone, on partial information.