Skip to content
MARL in Cooperative Environments
Edit this page

1.10Scaling Coordination

3 min read

Scaling coordination changes the size and difficulty of a cooperative problem even when each agent’s local action set remains fixed.

In this section you will

  • Derive the growth of the joint-action space
  • Identify the pressures scale puts on exploration and credit assignment
  • Explain what decentralized execution keeps manageable

The arithmetic is the same as ever.

Figure 1
∣A∣=∣A1∣×⋯×∣An∣\bigl|\mathcal{A}\bigr| = \bigl|\mathcal{A}_1\bigr| \times \cdots \times \bigl|\mathcal{A}_\nag\bigr|
∣A∣|\mathcal{A}|
the number of joint actions
∣Ai∣|\mathcal{A}_i|
the size of agent i’s local action space
n\nag
the number of agents
Each added agent multiplies the number of action combinations.
AgentsActions eachJoint actions
2636
461,296
861,679,616
166≈ 2.8 × 10¹²

Nothing subtle is happening. But the consequences reach further than the size of one table, and they compound.

Joint-action spaces grow rapidly. A method that must enumerate, search or represent joint actions is finished early. This rules out centralized action selection long before you reach an interesting number of agents.

Centralized value representations become harder. QtotQ_{\text{tot}} has to generalise over a space growing exponentially, from a number of samples growing nowhere near that fast. It is not only storage. It is that most of the space is never visited, so the value must be inferred rather than observed.

Non-stationarity increases. With two agents, each faces one moving partner. With n\nag agents, each faces n−1\nag - 1 of them, all updating simultaneously. The environment any one agent perceives drifts faster the more company it has.

Credit assignment becomes harder. One shared reward is consistent with more and more stories as the team grows. With two agents, a poor return has a handful of explanations; with sixteen, the same number is consistent with almost anything, and any individual agent’s share of it approaches noise.

Worth noting for balance, because it is the source of independent learning’s resilience.

An individual agent’s own problem does not grow. Agent i\ag still chooses among its own mm actions, from its own observation, updating its own function. That cost is flat in n\nag.

This is exactly why independent learning scales gracefully while centralized control does not, and why, on large teams, the comparison is often between independent learning and a factored CTDE method, with fully centralized approaches not in the running at all.

Knowledge check

A team grows from 3 agents to 12. Which effect is NOT a direct consequence of the joint action space growing?

Select one answer.

  • The joint action space is a product, so it grows exponentially in the number of agents.
  • Four things worsen together: joint-action size, the difficulty of representing a centralized value, non-stationarity (each agent now faces n−1\nag - 1 moving partners), and credit assignment.
  • Per-agent cost stays flat: each agent still chooses among its own actions from its own observation.
  • Decentralized action selection and factored values are therefore requirements at scale, not architectural taste.