Tutorial Structure
The tutorial follows one connected progression:
Background → Coordinate → Communicate → Adapt → LLMs as Cooperative Agents → Challenge Lab → Final Project
Expand each section to see the question it addresses and the main ideas covered.
Background: Cooperative MARL Foundations
Question: what changes when reinforcement learning moves from one agent to many?
- Reinforcement learning fundamentals
- Multi-agent environments
- States, observations, and actions
- Joint actions and policies
- Cooperative rewards and partial observability
- The Dec-POMDP framework
Chapter 1: Coordinate
Question: how can agents learn actions that work well together?
- Coordination and action interdependence
- Independent learning and non-stationarity
- Centralized and decentralized learning
- Centralized training with decentralized execution
- Credit assignment and centralized critics
- Value decomposition with VDN and QMIX
Chapter 2: Communicate
Question: what information should agents share when they know different things?
- Messages and communication actions
- Message content and communication policies
- Capacity, range, noise, loss, and cost
- Learned communication protocols
- Interpretable communication
- Communication failures and LangGround
Chapter 3: Adapt
Question: can an agent still cooperate when the other agents change?
- Partner dependence and partner generalization
- Training with diverse partners
- Agent modelling
- Partner representations
- Ad hoc teamwork and zero-shot coordination
- Cross-play and unseen-partner evaluation
Frontier: LLMs as Cooperative Agents
Question: what changes when the agents in a cooperative MARL system are large language models?
- LLMs as agents in interactive systems
- Multi-LLM collaboration
- Orchestrated versus learned collaboration
- Cooperative MARL and Dec-POMDP formulations
- Shared rewards, credit assignment, and CTDE for LLM agents
- Recent research, open problems, and emerging agent societies
Challenge Lab: Wireless Network Resource Allocation using MARL
Question: how can coordination, communication, and adaptation be applied to a new technical problem?
- Model wireless access points as cooperative agents
- Coordinate channel-selection decisions
- Reduce interference and improve throughput
- Add communication between access points
- Test changing traffic conditions
- Evaluate behaviour with an unfamiliar access point
Final Project: Design a MARL-Based Disaster-Response System
Question: how would you design and evaluate a complete cooperative MARL system?
- Define agent roles, observations, and actions
- Specify a shared objective
- Design a coordination approach
- Design a communication strategy
- Plan for unfamiliar agents and changing conditions
- Build an evaluation plan