Skip to content
MARL in Cooperative Environments
Edit this page

Learning LLM Collaboration with MARL

10 min read

The previous section mapped the parts of a multi-LLM system onto observations, policies, actions and rewards. This section takes that mapping seriously.

In this section you will

  • Write a multi-LLM system as a Dec-POMDP, term by term
  • Explain why a joint task reward reopens the credit-assignment problem
  • Describe how MAGRPO turns a group of sampled joint responses into a learning signal
  • Read what recent results establish, and what they do not
  • Name the questions this direction leaves open

We already have a framework for this.

Liu and colleagues define LLM collaboration with the same tuple the Background introduced:

⟨I, S, {Oi}, {Ai}, R, T, H⟩\left\langle \mathcal{I},\ \mathcal{S},\ \{\mathcal{O}_i\},\ \{\mathcal{A}_i\},\ R,\ T,\ H \right\rangle
  • I\mathcal{I} is the set of nn LLM agents, each instantiated with a pre-trained language model
  • S\mathcal{S} is the global state, split into what the system can see and the part of the user’s state it cannot
  • Oi\mathcal{O}_i holds agent ii‘s local observation, a natural-language instruction
  • Ai\mathcal{A}_i holds agent ii‘s local action, a response in natural language
  • RR is the joint reward, TT the stochastic transition, and HH the turn limit of the dialogue

Idea: Not one symbol of the framework changed. Only the contents of the observation and action spaces did.

Flip the switch below. The boxes do not move.

o¹π₁a¹o²π₂a²o³π₃a³Environmentshared reward

The same diagram in two vocabularies. An observation becomes a prompt, a policy becomes a language model, an action becomes a response, and the environment becomes the user or system being served.

Suppose two language models jointly produce a program. The finished program either passes its tests or it does not.

You may know that

Rteam=1R_{\text{team}} = 1

and still not know:

  • which agent’s response was most useful;
  • whether one response corrected another;
  • whether both were necessary;
  • which behaviour each policy should reinforce.

This is the credit assignment problem from Chapter 1, arriving unchanged in a new setting. One number came back for work that several agents did together, and it was never a sum of per-agent contributions, so it cannot be taken apart into them.

Liu and colleagues propose Multi-Agent Group Relative Policy Optimization, or MAGRPO, for multi-turn LLM collaboration. The idea is to judge a joint behaviour against other joint behaviours the same agents could have produced.

The method samples a group of joint responses, scores each with the shared objective, estimates a centralized group-relative advantage from those returns, and uses it to update the individual policies. The trained policies still execute in a decentralized way.

A^t(g)=Rt(g)−1G∑g′=1GRt(g′)σ(Rt(g))\hat{A}_t^{(g)} = \frac{R_t^{(g)} - \frac{1}{G}\sum_{g'=1}^{G} R_t^{(g')}}{\sigma\left(R_t^{(g)}\right)}
  • Rt(g)R_t^{(g)} is the return of the gg-th sampled joint response at turn tt
  • GG is the number of joint responses sampled as a group
  • σ\sigma is the standard deviation of the returns in that group

Idea: Compare each joint behaviour with the average of its own group, then divide by the spread so the size of the update does not depend on how the task happens to be scaled.

In four steps:

  1. generate several candidate joint behaviours;
  2. evaluate each with the shared objective;
  3. compare each return with the group average;
  4. reinforce the joint behaviours that did relatively well.

Group sampling and the group-relative comparison happen while learning. What is deployed is each policy on its own local input.

That is centralized training with decentralized execution, the paradigm from Chapter 1, in a setting where the actors are language models.

Visual Lab: Group-Relative Advantage Playground

Section titled “Visual Lab: Group-Relative Advantage Playground”

Four joint responses were sampled and scored. Move the sliders and watch which ones would be reinforced.

Four sampled joint responses

0.80
0.30
0.90
0.50
Group mean—
Standard deviation—
Reinforced—
CandidateScoreScore minus meanAdvantageEffect

Two behaviours are worth provoking deliberately:

  • Set every score the same. The mean carries no information, no candidate is preferred, and the learning signal vanishes. A group that never disagrees teaches nothing.
  • Add the same amount to every score. The mean moves with them and the advantages do not change. Only the ranking within the group matters, which is why this estimator does not need a separately learned value baseline.

Knowledge check

MAGRPO compares each sampled joint response with the average of its group. What does that comparison avoid having to do?

Select one answer.

Three papers, and what each one actually establishes.

  • Coordination. How should work be divided among LLM agents?
  • Communication. What information is actually worth exchanging, and how many turns are worth paying for?
  • Credit assignment. How should a joint outcome be attributed across agents and across turns?
  • Partner generalization. Will learned collaboration transfer to different models or different agent populations?
  • Scalability. How do you train several large policies without prohibitive sampling and inference cost? Liu and colleagues discuss this directly, noting the difficulty created by very large language action and observation spaces.
  • Evaluation. How do you distinguish genuine collaboration from simply making more model calls?

That last one should feel familiar. It is the question the Challenge Lab asked about communication: the messages improved throughput and still lowered the objective, because the objective charged for them. More calls is not more cooperation until something measures the difference.

  • Multi-LLM collaboration can be stated as a Dec-POMDP with no change to the framework: an observation is a prompt, an action is a whole response, and HH counts dialogue turns.
  • A joint task reward reopens credit assignment rather than resolving it.
  • MAGRPO samples a group of joint responses and scores each against the group’s own average, which avoids attributing the outcome to one agent.
  • Training uses group-level information; execution stays decentralized, which is CTDE in a new setting.
  • Current results are promising and narrow. Partner generalization, scalability, and honest evaluation are open.
  • Liu, S., Liang, Z., Lyu, X., and Amato, C. LLM Collaboration with Multi-Agent Reinforcement Learning. Proceedings of the AAAI Conference on Artificial Intelligence, 40(38), 32150 to 32158, 2026. Publisher · arXiv:2508.04652
  • Park, C., Han, S., Guo, X., Ozdaglar, A. E., Zhang, K., and Kim, J.-K. MAPoRL: Multi-Agent Post-Co-Training for Collaborative Large Language Models with Reinforcement Learning. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, 30215 to 30248, 2025. ACL Anthology
  • Liu, S., Chen, T., Amiri, R., and Amato, C. Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic. International Conference on Machine Learning, 2026. ICML · arXiv:2601.21972

Revisit the 5 key concepts, equations, and intuitions from this section.

Open Frontier Flashcards