Skip to content
MARL in Cooperative Environments
Edit this page

Parts 2 to 4: Coordinate, Communicate, Adapt

6 min read

The central project design combines three linked decisions. First, identify one coordination failure and select a training approach that addresses it. Second, specify useful message content, recipients, representation, and at least one communication constraint. Third, define a partner or environment change and choose a strategy for adaptation. Each choice must follow from the system formulation and state a behavioural consequence that the evaluation can later measure.


Name one, specifically. Examples of the right grain:

  • multiple drones searching the same area while another region goes uncovered,
  • two ground robots attempting the same rescue,
  • agents blocking one another in a corridor,
  • nobody maintaining the communication relay because every role is more urgent,
  • a poor division of roles when several agents could do several jobs.

“Coordination is hard” is not a coordination problem. Neither is a list of five. Pick the one that would most damage your system and describe how it arises from the observations and actions you defined in Part 1.

Then answer:

How should your training approach encourage coordinated behaviour?

Choose and justify one:

ApproachChoose it whenAccept that
Independent learningsimplicity and scale matter; you need a baselinenon-stationarity, high variance across seeds
Centralized criticlearning is unstable because partner behaviour confounds the signalone more component to train
Value decompositioncredit assignment is the bottleneck, and rewards are commonthe mixer grows with agent count
Centralized executionthe team is small and genuinely co-locatedan exponential joint action space, and no decentralized deployment

You do not have to use every technique in this resource. Saying which you rejected, and why, is worth more than using all of them.


Four decisions, and Part 1 should already have told you what the first one is.

Look back at the informational awkwardness you identified in Part 1: the decision an agent cannot make well from its own observation. What would fix it?

Candidates in this domain: discovered survivors, hazards, an intended route, task status, remaining battery.

Apply the three tests from Message Content to each:

  • Does the receiver already know it?
  • Can the receiver act on it?
  • Would it choose differently having heard it?

A message that fails any of the three is costing bandwidth for nothing.

Broadcast to all, nearby agents only, or specified recipients such as relay agents. Say which, and note that range makes communication depend on the state.

Discrete symbols, a small vector, or a structured status message. If discrete, say how many, because that is your capacity and it decides which situations must merge.

ConstraintWhat it forces you to think about
Bandwidthwhich distinctions to keep and which to merge
Message losswhether silence carries meaning in your protocol
Rangewhether the protocol survives agents moving apart
Costwhether each message is worth its price

Then answer:

What information is important enough to justify communication?


Now deployment uncertainty. Your system must handle at least one unfamiliar condition involving other agents.

  • a new rescue drone joins mid-operation,
  • one agent fails and stops responding,
  • robots from another organisation arrive,
  • an agent follows a policy you did not train,
  • the number of active agents changes.

The last two are the harder and more interesting ones. The first three can sometimes be handled by ordinary robustness.

How will the system avoid depending on one fixed group of agents?

Options, with what each costs:

StrategyCost
Diverse partner policiespeak performance with any one partner
Population-based trainingcompute, and the work of building the population
Randomised roles during trainingslower convergence
Behaviour variation across episodesyou must decide what to vary, and verify it varies

The warning from Training Partner Diversity applies: the number of partners is not behavioural diversity. Ten near-identical partners present one pattern. Say how you would check that yours actually differ.

Will your agent

  • infer other agents’ behaviour from observation,
  • maintain a partner representation,
  • react to observed actions without identifying anyone,
  • or rely on a shared communication convention?
Figure 1
observed behaviour⏟what I have seen→representation⏟who I think you are→my action⏟what I do about it\ubrace{observe}{\text{observed behaviour}}{what I have seen} \rightarrow \ubrace{policy}{\text{representation}}{who I think you are} \rightarrow \ubrace{action}{\text{my action}}{what I do about it}
observed behaviour\text{observed behaviour}
locally available evidence about another agent
representation\text{representation}
the inferred behavioural summary
my action\text{my action}
the policy response conditioned on that summary
An adaptation mechanism maps local behavioural evidence to a conditioned action.

This is the part most often skipped, and it is where the marks are.

How would the system behave differently after observing the unfamiliar agent?

An actual difference in action, not a statement that the system “adapts”. Something of this shape:

After three steps in which the new drone repeatedly enters sector B, my agents’ partner representation shifts toward area-focused, and their policy reallocates from sector B to sector D, leaving the new drone to finish B alone.


You should now have exactly:

  • one coordination problem, and one training approach that addresses it
  • one communication design: content, recipients, representation, constraint
  • one unfamiliar condition, a training strategy, an adaptation mechanism, and a concrete behavioural consequence

If you have several of each, cut. A design that commits is assessable; a survey is not.

Next: Part 5.