Skip to content
MARL in Cooperative Environments
Edit this page

Explainer Videos

Start here. It builds the subject from a single learning agent to cross-play evaluation, and each idea creates the problem that makes the next one necessary.

An Intro to Multi-Agent Reinforcement Learning in Cooperative Environments19:38Watch on YouTubeJoint actions, partial observability, shared rewards and credit assignment; CTDE, VDN and QMIX; learned communication under capacity, cost, delay and failure; partner dependence, ad hoc teamwork and cross-play. Ends on four questions the field has not settled.

The frontier chapter, in video form. It carries the same ideas to teams of language models and is careful about which claims current research supports.

When LLM Agents Work Together: Cooperative Multi-Agent Reinforcement Learning Meets Language Models13:32Watch on YouTubeDesigned versus learned collaboration; multi-LLM collaboration as a Dec-POMDP; why a shared reward over a joint response reopens credit assignment; MAGRPO and group-relative advantage; agent societies and what remains unsolved.
Video sectionWritten chapters
Joint actions, observations, shared rewardBackground
Independent learning, CTDE, credit assignment, VDN, QMIXCoordinate
Message capacity, cost, failure, learned protocolsCommunicate
Partner dependence, diversity, ad hoc teamwork, cross-playAdapt
Everything in the second videoLLMs as Cooperative Agents