1.10Scaling Coordination
Scaling coordination changes the size and difficulty of a cooperative problem even when each agent’s local action set remains fixed.
In this section you will
- Derive the growth of the joint-action space
- Identify the pressures scale puts on exploration and credit assignment
- Explain what decentralized execution keeps manageable
Joint-Action Space Growth
Section titled “Joint-Action Space Growth”The arithmetic is the same as ever.
- the number of joint actions
- the size of agent i’s local action space
- the number of agents
| Agents | Actions each | Joint actions |
|---|---|---|
| 2 | 6 | 36 |
| 4 | 6 | 1,296 |
| 8 | 6 | 1,679,616 |
| 16 | 6 | ≈ 2.8 × 10¹² |
Nothing subtle is happening. But the consequences reach further than the size of one table, and they compound.
Scaling Pressures
Section titled “Scaling Pressures”Joint-action spaces grow rapidly. A method that must enumerate, search or represent joint actions is finished early. This rules out centralized action selection long before you reach an interesting number of agents.
Centralized value representations become harder. has to generalise over a space growing exponentially, from a number of samples growing nowhere near that fast. It is not only storage. It is that most of the space is never visited, so the value must be inferred rather than observed.
Non-stationarity increases. With two agents, each faces one moving partner. With agents, each faces of them, all updating simultaneously. The environment any one agent perceives drifts faster the more company it has.
Credit assignment becomes harder. One shared reward is consistent with more and more stories as the team grows. With two agents, a poor return has a handful of explanations; with sixteen, the same number is consistent with almost anything, and any individual agent’s share of it approaches noise.
Benefits of Decentralized Execution
Section titled “Benefits of Decentralized Execution”Worth noting for balance, because it is the source of independent learning’s resilience.
An individual agent’s own problem does not grow. Agent still chooses among its own actions, from its own observation, updating its own function. That cost is flat in .
This is exactly why independent learning scales gracefully while centralized control does not, and why, on large teams, the comparison is often between independent learning and a factored CTDE method, with fully centralized approaches not in the running at all.
Knowledge check
Correct.
Not quite.
Each agent’s own action selection becomes more expensive.
Correct. This is the thing that does not happen. Each agent still chooses among its own m actions from its own observation, so per-agent decision cost is flat in the number of agents. That is the property decentralized methods rely on.
A centralized value must generalise over exponentially more joint actions.
This is a direct consequence. The space grows exponentially while the sample count does not, so most of it must be inferred rather than observed.
A shared reward becomes consistent with far more explanations.
This is a direct consequence, and it is the credit assignment problem worsening. More agents means more combinations that could have produced the same team return.
Enumerating joint actions to check a policy becomes infeasible.
This is a direct consequence. Exhaustive enumeration is available only for small joint-action spaces, so larger problems require structured values or sampling.
Explanation
Knowing which costs grow and which stay flat is what tells you where a method will break.
Scaling Coordination Summary
Section titled “Scaling Coordination Summary”- The joint action space is a product, so it grows exponentially in the number of agents.
- Four things worsen together: joint-action size, the difficulty of representing a centralized value, non-stationarity (each agent now faces moving partners), and credit assignment.
- Per-agent cost stays flat: each agent still chooses among its own actions from its own observation.
- Decentralized action selection and factored values are therefore requirements at scale, not architectural taste.