Skip to content
MARL in Cooperative Environments
Edit this page

Part 1: Define the System

4 min read

A defensible MARL design begins with a precise system formulation. This section guides you through selecting agent roles, defining the complete environment state, assigning local observations and actions, constructing one shared reward, and stating the team objective. Each later coordination, communication, and adaptation choice will be evaluated against these definitions. The goal is a Dec-POMDP specification that preserves meaningful decentralized information constraints.

Pick three to five roles. Not fifteen.

For each, specify what it sees and what it can do:

RoleObservationActions
Search dronelocal map tile, detected people, battery, nearby agentsmove, scan, return to base, relay
your role 2
your role 3

That first row is an example, not a requirement. Replace it.

What would the full environment state contain, whether or not anybody can see it?

This matters even though no agent receives it, for two reasons. It is what your transition and reward functions are defined on. And writing it down is what lets you check that your observations are genuinely smaller than it.

A useful test from the Background: is anything in your state actually a belief? If your state includes “the team’s estimate of where survivors are”, that is not state, it is inference.

For each role, what does that agent actually receive each step?

Figure 1
oti=Oi(st)\tone{observe}{\obs{\ag}} = O_\ag\bigl(\tone{observe}{\st}\bigr)
OiO_\ag
this role’s observation function: the crop of the truth it gets
An observation is a filter on the state. Each role has its own.

Two properties to check explicitly, because Part 3 depends on them:

Asymmetry. Do different roles hold different information? If every agent sees the same thing, there is nothing to communicate and Part 3 will be hollow.

Insufficiency. Is there a decision an agent must make where its own observation is not enough? That is your communication or inference requirement, and naming it here makes Parts 3 and 4 straightforward.

What decisions can each role make? Keep the sets small and discrete unless you have a reason not to.

One constraint worth stating, because it catches people: no agent may act on another agent. “Assign drone 2 to sector B” is not an action unless you are designing a centralized controller, in which case say so and accept the consequences from the Coordinate chapter. Otherwise an agent acts, and the consequences reach its partners through the environment or through a message.

One team reward. Every agent receives the same number.

Keep it to three or four terms. The Challenge Lab used three deliberately.

Figure 2
rt=w1⋅people assisted⏟what you want−w2⋅response time⏟what you want less of−w3⋅unsafe actions⏟what you will not accept\tone{reward}{r_t} = \ubrace{reward}{w_1 \cdot \text{people assisted}}{what you want} - \ubrace{conflict}{w_2 \cdot \text{response time}}{what you want less of} - \ubrace{conflict}{w_3 \cdot \text{unsafe actions}}{what you will not accept}
rtr_t
the shared disaster-response reward
w1,w2,w3w_1,w_2,w_3
weights expressing the relative design priorities
unsafe actions\text{unsafe actions}
a safety penalty shared by the whole team
An example reward balances assistance, response time, and safety.

For each term, be ready to say what behaviour it is there to produce, and what it might accidentally reward instead.

One sentence, in plain words: what is the system collectively trying to achieve?

Then check it against your reward. If the sentence and the reward function disagree, the reward wins, and that is a bug rather than a philosophical position.

  • Three to five roles, each with observations and actions
  • A state that is genuinely larger than any observation
  • At least one role holding information another role needs
  • At least one decision an agent cannot make well from its own observation
  • One shared reward, three or four terms, each justified
  • One sentence stating the collective objective

That last-but-one item is the hinge. If nothing in your system is informationally awkward, Parts 3 and 4 will have nothing to work on, and it is worth going back and making the observations narrower.

Next: Parts 2 to 4.