Skip to content
MARL in Cooperative Environments
Edit this page

Final Project Overview

3 min read

The final project asks you to design a cooperative multi-agent system for disaster response. You will define heterogeneous agent roles, states, local observations, actions, and a shared reward; select coordination, communication, and adaptation mechanisms; and propose an evaluation that can expose failure under realistic variation. No algorithm or environment is prescribed. By the end, your design should make clear how decentralized agents cooperate, what assumptions support the chosen methods, and which evidence would justify deployment claims.

Figure 1
Learn→Experiment→Transfer→Create\text{Learn} \rightarrow \text{Experiment} \rightarrow \text{Transfer} \rightarrow \cbox{policy}{\text{Create}}
Learn\text{Learn}
develop the conceptual model
Experiment\text{Experiment}
test mechanisms in chapter labs
Transfer\text{Transfer}
apply the concepts in a supplied new domain
Create\text{Create}
design and evaluate a new cooperative system
The final project applies the course concepts to a system designed by the learner.

A natural disaster has affected a region. A multi-agent system is deployed to

  • locate people,
  • assess hazards,
  • deliver supplies,
  • maintain communication,
  • coordinate rescue activities.

The system may include aerial drones, ground robots, communication relay agents and supply-delivery robots.

Three things are true of it, and each one is a chapter:

No single agent has complete information. The region is large, sensors are local, and the map is out of date.

Communication can be unreliable. Infrastructure is damaged, range is limited, and bandwidth is contested.

Some agents arrive later, or come from organisations that were not part of the original system. You did not train them, cannot inspect them, and cannot change them.

How would you design a cooperative multi-agent system that can coordinate, communicate, and adapt during disaster response?

PartWhat you produce
1 · Define the systemagents, state, observations, actions, reward
2 · Coordinateone coordination problem, and a training approach that addresses it
3 · Communicatewhat is sent, to whom, in what representation, under what constraint
4 · Adapthow the system handles agents it did not train with
5 · Evaluatefour to six metrics, and the test conditions that would distinguish a robust system from a lucky one

Parts 2 to 4 are on one page, because they are the same design decision seen from three angles.

Three artifacts, and they are deliberately lean:

  1. One system diagram.
  2. A design brief, maximum two pages.
  3. An evaluation plan, one compact table.

No long essay. The Deliverables page has the detail, and there is an optional implementation extension for anyone who wants it.

Start with Part 1.