1.6Centralized Critics
A centralized critic is one way to use global information under CTDE: the critic evaluates actions during training while each actor remains a local policy for execution.
In this section you will
- Separate the roles of the actor and the critic
- Expand the critic’s inputs to the state and the joint action
- Explain how a richer evaluation changes the policy update
- State what remains decentralized at execution
Actor and Critic Roles
Section titled “Actor and Critic Roles”Many reinforcement learning methods split the learner in two.
The actor is the policy: the thing that chooses actions. The critic is a learned value or advantage function used to improve that policy. It does not choose anything, it evaluates.
- the actor: the policy, with its own parameters θ
- the critic: a learned estimate of how good a situation is, parameters φ
The critic exists to make the actor’s updates less noisy. Instead of judging an action by the whole messy return that followed it, the actor is nudged according to how that action compared with what the critic expected.
And here is the observation that makes this section work: only the actor is needed at deployment. The critic is training scaffolding. Once the policy is trained, the critic is discarded.
Centralized Critic Information
Section titled “Centralized Critic Information”That is the whole idea. Condition the critic on the centralized information , and leave the actor exactly as constrained as it has to be.
- agent i’s own history, what it has personally seen and done
- centralized information, available because this is training
- the joint action; an action-value critic is conditioned on what everybody did
Both forms appear in the literature. A centralized state-value critic is conditioned on the agent’s history plus the centralized information. A centralized action-value critic is additionally conditioned on the joint action, which, from Coordination, is exactly the form that can tell combinations apart.
Meanwhile the actor is untouched:
- the actor: still local, still deployable, still the only thing that runs
- its own history and nothing else, no z, no partner action
The picture is the CTDE picture with one box renamed, and that is the point. A centralized critic is not a different arrangement. It is a specific answer to “what consumes ?”
Effects on the Learning Signal
Section titled “Effects on the Learning Signal”The critic’s job is to say how good things were. A blind critic has to do that from one agent’s partial view, which means the same observation gets the same evaluation whether the team’s other half was doing something useful or something disastrous. Its estimates are noisy for a reason that has nothing to do with the agent being evaluated.
Give it and it can see what the team actually did. Its estimates get sharper, the actor’s updates get less noisy, and learning gets more stable, all without changing what the actor is allowed to know.
Knowledge check
Correct.
Not quite.
The critic is discarded after training, so it never has to run under the deployment constraint.
Exactly. The decentralization constraint applies to what executes. The critic exists only to shape the actor’s updates, so conditioning it on anything training can provide costs nothing at deployment.
The critic only estimates values, and values are less sensitive than actions.
Nothing about values is inherently less sensitive. The reason is about lifetime, not about the type of quantity: the critic does not exist at deployment, so its inputs are unconstrained.
The actor indirectly receives the global state through the critic’s gradients, so both effectively see it.
The first half is true and the conclusion is wrong. Information does reach the actor’s parameters through the gradients. That is the mechanism. But the deployed actor still cannot condition on the state it does not receive, so it has not "effectively seen" it in any usable sense.
Because the global state is usually similar to the agent’s observation anyway.
Often false, and if it were true the centralized critic would be pointless. The value comes precisely from the state containing what the observation lacks.
Explanation
“Which components have to exist at deployment?” is the question that resolves most CTDE design arguments.
Centralized Critics Summary
Section titled “Centralized Critics Summary”- An actor chooses actions; a critic evaluates. Only the actor is needed at deployment.
- A centralized critic is conditioned on the agent’s history plus centralized information ; an action-value version is additionally conditioned on the joint action .
- The actor stays local: .
- The benefit is a better-informed learning signal, the critic can see what the team did, so its estimates are less confounded by partner behaviour.
- It does not solve coordination. The actor still decides alone, on partial information.