How to Use This Tutorial
Use this page as your map from cooperative MARL foundations to current research frontiers, practical application, and system design.
Your Learning Route
Section titled “Your Learning Route”1. Build the Foundations
Section titled “1. Build the Foundations”Start here if reinforcement learning or MARL is new to you.
You will build the foundation for:
- multiple agents;
- states and observations;
- joint actions and policies;
- cooperative rewards;
- partial observability;
- Dec-POMDPs.
↓
2. Chapter 1: Coordinate
Section titled “2. Chapter 1: Coordinate”Question: How should the agents’ actions fit together?
Follow:
Introduction → Concepts → Knowledge Checks → Interactive Virtual Lab → Worksheet
You will:
- understand the coordination problem;
- compare different training approaches;
- experiment with coordination directly in the website;
- consolidate the main ideas and trade-offs.
↓
3. Chapter 2: Communicate
Section titled “3. Chapter 2: Communicate”Question: What information should agents share?
Follow:
Introduction → Concepts → Knowledge Checks → Colab Lab → Worksheet
You will:
- understand why communication can help;
- study communication constraints and protocols;
- train agents to learn a communication protocol;
- test capacity, noise, and protocol mismatch.
↓
4. Chapter 3: Adapt
Section titled “4. Chapter 3: Adapt”Question: What happens when the other agents change?
Follow:
Introduction → Concepts → Knowledge Checks → Colab Lab → Worksheet
You will:
- distinguish familiar-partner performance from general cooperation;
- study partner diversity and agent modelling;
- evaluate agents with unseen partners;
- observe adaptation when partner behaviour changes.
↓
5. Explore the Frontier
Section titled “5. Explore the Frontier”Question: What changes when the cooperating agents are large language models?
Use what you learned about Coordinate + Communicate + Adapt to explore a rapidly developing research direction.
You will:
- map LLM collaboration to the cooperative MARL framework;
- compare orchestrated and learned collaboration;
- connect shared rewards, credit assignment, and CTDE to multi-LLM training;
- explore recent research on learning collaboration between LLM agents;
- consider open questions around generalization, scalability, and agent societies.
↓
6. Apply Everything
Section titled “6. Apply Everything”Challenge Lab: Wireless Network Resource Allocation using MARL
Section titled “Challenge Lab: Wireless Network Resource Allocation using MARL”Apply:
Coordinate + Communicate + Adapt
You will work with wireless access points that must:
- share a limited number of channels;
- reduce interference;
- coordinate using local information;
- communicate when useful;
- remain effective under changing conditions.
↓
7. Create Your Own System
Section titled “7. Create Your Own System”Final Project: Design a MARL-Based Disaster-Response System
Section titled “Final Project: Design a MARL-Based Disaster-Response System”Bring everything together.
You will decide:
- what the agents are;
- what they observe;
- what actions they can take;
- how they coordinate;
- what they communicate;
- how they adapt;
- how the system should be evaluated.
Choose Your Starting Point
Section titled “Choose Your Starting Point”- New to reinforcement learning: start with Background.
- Know basic RL: begin with From One Agent to Many.
- Know MARL foundations: begin with Chapter 1: Coordinate.
- Reviewing one topic: jump directly to the relevant chapter.
Within Each Chapter
Section titled “Within Each Chapter”Use this simple rhythm:
Understand → Test → Experiment → Consolidate
- Understand: work through the concepts and visual explanations.
- Test: complete the embedded knowledge checks.
- Experiment: use the practical lab.
- Consolidate: complete the worksheet.