Learning LLM Collaboration with MARL
The previous section mapped the parts of a multi-LLM system onto observations, policies, actions and rewards. This section takes that mapping seriously.
In this section you will
- Write a multi-LLM system as a Dec-POMDP, term by term
- Explain why a joint task reward reopens the credit-assignment problem
- Describe how MAGRPO turns a group of sampled joint responses into a learning signal
- Read what recent results establish, and what they do not
- Name the questions this direction leaves open
1. The Dec-POMDP Connection
Section titled “1. The Dec-POMDP Connection”We already have a framework for this.
Liu and colleagues define LLM collaboration with the same tuple the Background introduced:
Key Equation: The Dec-POMDP, Unchanged
Section titled “Key Equation: The Dec-POMDP, Unchanged”
- is the set of LLM agents, each instantiated with a pre-trained language model
- is the global state, split into what the system can see and the part of the user’s state it cannot
- holds agent ‘s local observation, a natural-language instruction
- holds agent ‘s local action, a response in natural language
- is the joint reward, the stochastic transition, and the turn limit of the dialogue
Idea: Not one symbol of the framework changed. Only the contents of the observation and action spaces did.
Flip the switch below. The boxes do not move.
The same diagram in two vocabularies. An observation becomes a prompt, a policy becomes a language model, an action becomes a response, and the environment becomes the user or system being served.
2. Why a Shared Reward Matters
Section titled “2. Why a Shared Reward Matters”Suppose two language models jointly produce a program. The finished program either passes its tests or it does not.
You may know that
and still not know:
- which agent’s response was most useful;
- whether one response corrected another;
- whether both were necessary;
- which behaviour each policy should reinforce.
This is the credit assignment problem from Chapter 1, arriving unchanged in a new setting. One number came back for work that several agents did together, and it was never a sum of per-agent contributions, so it cannot be taken apart into them.
3. MAGRPO: One Recent Approach
Section titled “3. MAGRPO: One Recent Approach”Liu and colleagues propose Multi-Agent Group Relative Policy Optimization, or MAGRPO, for multi-turn LLM collaboration. The idea is to judge a joint behaviour against other joint behaviours the same agents could have produced.
The method samples a group of joint responses, scores each with the shared objective, estimates a centralized group-relative advantage from those returns, and uses it to update the individual policies. The trained policies still execute in a decentralized way.
Key Equation: Group-Relative Advantage
Section titled “Key Equation: Group-Relative Advantage”
- is the return of the -th sampled joint response at turn
- is the number of joint responses sampled as a group
- is the standard deviation of the returns in that group
Idea: Compare each joint behaviour with the average of its own group, then divide by the spread so the size of the update does not depend on how the task happens to be scaled.
In four steps:
- generate several candidate joint behaviours;
- evaluate each with the shared objective;
- compare each return with the group average;
- reinforce the joint behaviours that did relatively well.
Training and execution are separate
Section titled “Training and execution are separate”Group sampling and the group-relative comparison happen while learning. What is deployed is each policy on its own local input.
That is centralized training with decentralized execution, the paradigm from Chapter 1, in a setting where the actors are language models.
Visual Lab: Group-Relative Advantage Playground
Section titled “Visual Lab: Group-Relative Advantage Playground”Four joint responses were sampled and scored. Move the sliders and watch which ones would be reinforced.
Four sampled joint responses
| Candidate | Score | Score minus mean | Advantage | Effect |
|---|
Two behaviours are worth provoking deliberately:
- Set every score the same. The mean carries no information, no candidate is preferred, and the learning signal vanishes. A group that never disagrees teaches nothing.
- Add the same amount to every score. The mean moves with them and the advantages do not change. Only the ranking within the group matters, which is why this estimator does not need a separately learned value baseline.
Knowledge check
Correct.
Not quite.
Decide which individual agent was responsible for the outcome.
Right. The comparison is between whole joint behaviours, so the update never has to attribute the score to one agent. That is a way of proceeding despite the credit-assignment problem rather than a solution to it, and the distinction is worth keeping.
Collect a reward at all, since the group average replaces it.
The group average is computed from the rewards. Every candidate still has to be scored by the shared objective; what changes is what each score is compared against.
Train a separate value function as a baseline.
This is true and it is a real convenience, but it is not what the comparison avoids in the sense the question asks. Using the group mean as the baseline is a consequence of sampling a group, not the reason for doing so.
Keep execution decentralized.
The opposite: decentralized execution is preserved deliberately, and the group-relative comparison is part of the centralized training that makes it possible. That is the CTDE pattern from Chapter 1.
Explanation
A method that works despite a problem and a method that solves it are different claims, and they support different conclusions.
4. What Current Research Shows
Section titled “4. What Current Research Shows”Three papers, and what each one actually establishes.
5. What Remains Open
Section titled “5. What Remains Open”- Coordination. How should work be divided among LLM agents?
- Communication. What information is actually worth exchanging, and how many turns are worth paying for?
- Credit assignment. How should a joint outcome be attributed across agents and across turns?
- Partner generalization. Will learned collaboration transfer to different models or different agent populations?
- Scalability. How do you train several large policies without prohibitive sampling and inference cost? Liu and colleagues discuss this directly, noting the difficulty created by very large language action and observation spaces.
- Evaluation. How do you distinguish genuine collaboration from simply making more model calls?
That last one should feel familiar. It is the question the Challenge Lab asked about communication: the messages improved throughput and still lowered the objective, because the objective charged for them. More calls is not more cooperation until something measures the difference.
Learning LLM Collaboration Summary
Section titled “Learning LLM Collaboration Summary”- Multi-LLM collaboration can be stated as a Dec-POMDP with no change to the framework: an observation is a prompt, an action is a whole response, and counts dialogue turns.
- A joint task reward reopens credit assignment rather than resolving it.
- MAGRPO samples a group of joint responses and scores each against the group’s own average, which avoids attributing the outcome to one agent.
- Training uses group-level information; execution stays decentralized, which is CTDE in a new setting.
- Current results are promising and narrow. Partner generalization, scalability, and honest evaluation are open.
References
Section titled “References”- Liu, S., Liang, Z., Lyu, X., and Amato, C. LLM Collaboration with Multi-Agent Reinforcement Learning. Proceedings of the AAAI Conference on Artificial Intelligence, 40(38), 32150 to 32158, 2026. Publisher · arXiv:2508.04652
- Park, C., Han, S., Guo, X., Ozdaglar, A. E., Zhang, K., and Kim, J.-K. MAPoRL: Multi-Agent Post-Co-Training for Collaborative Large Language Models with Reinforcement Learning. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, 30215 to 30248, 2025. ACL Anthology
- Liu, S., Chen, T., Amiri, R., and Amato, C. Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic. International Conference on Machine Learning, 2026. ICML · arXiv:2601.21972
Review with Flashcards
Section titled “Review with Flashcards”Revisit the 5 key concepts, equations, and intuitions from this section.