Interactive learning resource
Multi-Agent Reinforcement Learning in Cooperative Environments
Coordinate, Communicate, Adapt
In multi-agent reinforcement learning, where multiple agents learn to act toward a shared goal, generalizable cooperation asks whether effective teamwork can persist when partners, communication conventions, or deployment conditions differ from training.
Your Learning Journey
A problem-driven path, not an algorithm survey. Each stage creates the difficulty that makes the next one necessary, and techniques arrive as tools for solving it.

- 01
Coordinate
Learn how agents align their actions toward a common goal.
- Read & ExploreInteractive content and examples
- Knowledge ChecksTest your reasoning as you go
- Lab 01Hands-on experiments
- 02
Communicate
Share information to overcome limits in what each agent can see.
- Read & ExploreInteractive content and examples
- Knowledge ChecksTest your reasoning as you go
- Lab 02Hands-on experiments
- 03
Adapt
Generalize to new teammates and situations, not just familiar ones.
- Read & ExploreInteractive content and examples
- Knowledge ChecksTest your reasoning as you go
- Lab 03Hands-on experiments
Where the three chapters land
LLMs as Cooperative Agents
Several language models working on one task are a cooperative multi-agent system, and every question from the three chapters returns: they act on partial information, they are paid on a shared outcome, and nothing tells an individual model what its own contribution was worth. This chapter reads the current research through the vocabulary you have just built, and is careful about which claims the evidence supports.
Read the Frontier Chapter- Prompted or learned the difference between arranging LLMs in a conversation and training them to collaborate
- The same formalism multi-LLM collaboration written as a Dec-POMDP, with the same partial observability
- Credit, again why one shared reward over a joint response reopens the problem from Coordinate
- What is shown MAGRPO and the current results, separated from what they do not yet establish
Apply it to the real world
Challenge Lab
Wireless Network Resource Allocation
Several base stations serve users over a limited set of channels. Each station chooses a channel and a transmit power using only what it can measure locally. Choose well and the network carries more traffic; choose the same channel as your neighbour and you both lose.
Go to Challenge LabCreate something of your own
Final Project
Design and justify a cooperative resource-allocation system: choose the observations, the actions, the reward, the communication scheme and the training conditions, then defend your deployment decision with measured evidence.
Start Final Project- Apply build on the lab environment
- Analyse seven evaluation scenarios
- Evaluate throughput against fairness
- Create propose an improved system
Resources
Everything you need to learn, practise, and explore cooperative multi-agent reinforcement learning.
- Interactive Tutorial WebsiteLearn each concept through clear explanations, visual examples, interactive demonstrations, and real-world connections.
- Explainer VideosTwo narrated animated explainers: the core cooperative MARL journey, and the frontier of LLMs as cooperative agents.
- Hands-On Practical NotebooksImplement key ideas, run experiments, analyse results, and explore how multi-agent systems behave.
- Challenge LabTransfer all three ideas to a wireless network: coordinate channels, price communication, and meet an access point you did not train.
- Final ProjectDesign a cooperative multi-agent disaster response system, and the evaluation that could prove it does not work.
- WorksheetsGuided exercises and activities for working through concepts, experiments, and problems at your own pace.
- Knowledge ChecksShort questions embedded throughout the course to test your understanding and provide immediate feedback.
- FlashcardsForty-five cards across five decks, laid out as a study sheet. Turn any card over, search the set, or filter to one idea.
- Python Packagecooperative-marl-labs: the environments, agents, training loops and evaluation used throughout, installable and documented on its own.