Introduction
Many real-world problems involve more than one decision-maker.
- robots may need to work together;
- autonomous vehicles share roads with other vehicles;
- wireless access points share limited communication channels;
- disaster-response agents may need to divide tasks and exchange information;
- AI agents, including large language models, may need to coordinate, exchange information, and work with unfamiliar agents.
In these settings, being individually intelligent is not enough.
An agent’s success can depend on:
- what the other agents do;
- what information they know;
- how their actions interact;
- whether the agents around them behave as expected.
This is the focus of Multi-Agent Reinforcement Learning (MARL). In this tutorial, we use those foundations to build toward one emerging question:
How can agents cooperate when the partners, protocols, or conditions around them change?
Coordinate
actions that fit together
Communicate
information that reaches the other agent
Adapt
a partner you did not train with
The three capabilities this tutorial develops: agents coordinating actions, communicating information, and adapting to a partner they did not train with.
Why This Matters
Section titled “Why This Matters”Two individually capable agents can still fail when working together.
For example:
- both agents may choose the same task when different tasks are needed;
- one agent may know something the other needs;
- a strategy that works with one partner may fail with another;
- individually reasonable actions may combine into a poor team outcome.
Cooperative MARL therefore asks a broader question than single-agent reinforcement learning:
How can multiple agents learn to work together toward a shared goal?
We will approach that question through three connected capabilities, then use them to study partner generalization, zero-shot coordination, and LLM-agent collaboration.
Chapter 1: Coordinate
Section titled “Chapter 1: Coordinate”How can agents learn actions that work well together?
You will explore:
- joint decisions;
- independent learning;
- centralized training;
- credit assignment;
- value decomposition.
Chapter 2: Communicate
Section titled “Chapter 2: Communicate”What information should agents share when they know different things?
You will explore:
- messages;
- communication constraints;
- communication policies;
- learned protocols;
- interpretable communication.
Chapter 3: Adapt
Section titled “Chapter 3: Adapt”Can agents still cooperate when the other agents change?
You will explore:
- partner dependence;
- partner diversity;
- agent modelling;
- ad hoc teamwork;
- zero-shot coordination.
Exploring the Frontier
Section titled “Exploring the Frontier”The three chapters build a foundation for a rapidly developing question in AI:
What happens when the cooperating agents are large language models?
Recent research is beginning to treat teams of LLMs as cooperative learning systems rather than only as independently prompted assistants.
In LLMs as Cooperative Agents, you will connect the ideas developed throughout the tutorial to this emerging frontier:
- modelling LLMs as agents;
- distinguishing orchestrated collaboration from learned collaboration;
- formulating multi-LLM collaboration as a cooperative MARL problem;
- connecting shared rewards, credit assignment, and CTDE to LLM training;
- exploring current research and open questions around increasingly complex agent teams and societies.
The goal is not simply to learn established MARL methods. It is to build the foundation needed to understand where cooperative learning is going next.
The Learning Journey
Section titled “The Learning Journey”Background → Coordinate → Communicate → Adapt → Explore the Frontier
The core chapters progressively build the ideas needed to reason about cooperative agents:
- Coordinate actions with other agents.
- Communicate when useful information is distributed.
- Adapt when the agents around you are unfamiliar or changing.
- Explore the Frontier by connecting these ideas to emerging research on LLMs as cooperative agents.
You will then apply and extend what you have learned through:
- Challenge Lab: Wireless Network Resource Allocation using MARL
- Final Project: Design a MARL-Based Disaster-Response System
By the End
Section titled “By the End”You should be able to:
- formulate a cooperative multi-agent problem;
- reason about coordination, communication, and adaptation;
- compare different MARL approaches and their assumptions;
- evaluate cooperation under familiar and unfamiliar conditions;
- connect cooperative MARL concepts to emerging research on LLM collaboration;
- design and justify a cooperative MARL system.