Skip to content
MARL in Cooperative Environments
Edit this page

Learning Objectives

The goal of this tutorial is not only to recognize MARL terminology.

You will progress from understanding cooperative MARL concepts to applying and analysing them, connecting them to emerging research, evaluating multi-agent systems, and designing your own cooperative system.

The tutorial progresses through six levels:

  • Remember: recognize the main components and terminology of cooperative MARL.
  • Understand: explain why coordination, communication, and adaptation are difficult.
  • Apply: use MARL concepts in interactive and computational experiments.
  • Analyse: compare behaviours, methods, communication protocols, and partner interactions.
  • Evaluate: test systems under changing conditions and justify design decisions.
  • Create: design a complete cooperative MARL system.

Bloom’s levels: Remember → Understand

By the end of the Background section, you should be able to:

  • distinguish states, observations, individual actions, and joint actions;
  • explain how individual policies combine into a joint policy;
  • explain shared rewards and partial observability;
  • describe why multiple learning agents create new challenges;
  • formulate a cooperative task using the Dec-POMDP framework.

Bloom’s levels: Understand → Apply → Analyse

By the end of the Coordinate chapter, you should be able to:

  • explain why shared rewards do not automatically produce coordinated behaviour;
  • identify action interdependence and multi-agent non-stationarity;
  • compare independent, centralized, decentralized, and CTDE approaches;
  • explain centralized critics and the credit-assignment problem;
  • explain how VDN and QMIX connect individual utilities to team value;
  • analyse learned coordination using the interactive virtual lab.

Bloom’s levels: Understand → Apply → Analyse

By the end of the Communicate chapter, you should be able to:

  • represent communication as part of an agent’s action;
  • identify information that may be useful to communicate;
  • reason about message capacity, cost, range, noise, and loss;
  • compare hand-designed and learned communication protocols;
  • inspect and interpret an emergent communication protocol;
  • analyse why communication may fail when agents use different conventions.

Bloom’s levels: Understand → Apply → Analyse

By the end of the Adapt chapter, you should be able to:

  • distinguish partner specialization from general cooperation;
  • explain why training with diverse partners can improve robustness;
  • use simple agent modelling to infer another agent’s behaviour;
  • explain ad hoc teamwork and zero-shot coordination;
  • analyse familiar-partner and unseen-partner performance;
  • interpret cross-play and partner-generalization results.

Bloom’s levels: Understand → Analyse

By the end of the LLMs as Cooperative Agents section, you should be able to:

  • explain how an LLM can be modelled as an agent in an interactive system;
  • map observations, policies, actions, rewards, and environment transitions to multi-LLM collaboration;
  • distinguish prompted or orchestrated collaboration from learned collaboration;
  • explain why coordination, communication, credit assignment, and adaptation remain relevant when the agents are LLMs;
  • connect CTDE and shared team rewards to recent approaches for training collaborating LLM agents;
  • identify open research questions in multi-LLM cooperation and larger agent societies.

Challenge Lab: Wireless Network Resource Allocation using MARL

Section titled “Challenge Lab: Wireless Network Resource Allocation using MARL”

Bloom’s levels: Apply → Analyse → Evaluate

By the end of the Challenge Lab, you should be able to:

  • formulate wireless access points as cooperative agents;
  • apply coordination methods to channel-selection decisions;
  • investigate when communication improves network performance;
  • test policies under changing traffic and unfamiliar-agent behaviour;
  • compare alternative MARL system designs;
  • justify a deployment decision using experimental evidence.

Final Project: Design a MARL-Based Disaster-Response System

Section titled “Final Project: Design a MARL-Based Disaster-Response System”

Bloom’s level: Create

By the end of the Final Project, you should be able to:

  • formulate a cooperative disaster-response problem;
  • define agent roles, observations, actions, and shared objectives;
  • select and justify a coordination approach;
  • design a communication strategy;
  • design for unfamiliar agents and changing conditions;
  • propose an evaluation that tests the system’s claims.

Remember → Understand → Apply → Analyse → Evaluate → Create

Across the tutorial:

Background → Coordinate → Communicate → Adapt → Explore the Frontier → Apply → Create