Skip to content
MARL in Cooperative Environments
Edit this page

2.9Research Connection: LangGround

6 min read

LangGround illustrates how learned communication can remain useful while its semantics are anchored to human language.

In this section you will

  • State the transfer problem created by private emergent conventions
  • Describe how language grounding supplies an external reference space
  • Report what the authors measured, and on which axes
  • Extract the design principle that generalizes beyond the paper

Language Grounded Multi-agent Reinforcement Learning with Human-interpretable Communication (NeurIPS 2024) opens by noting that MARL methods can enable agents to learn a shared communication protocol from scratch and accomplish challenging team tasks, but that the learned language is usually not interpretable to humans or to other agents not co-trained together, which limits its applicability in ad hoc teamwork.

That is the two-part finding from Interpretable Communication, arrived at independently: opacity to people, and incompatibility with agents you did not train with. The second is the one with teeth, and it is why this paper belongs at the end of a communication chapter rather than in a section on explainability.

The idea, at the level worth carrying:

Rather than leaving the communication representation entirely unconstrained, align the agents’ communication space with an embedding space derived from human natural language.

Concretely, the authors ground agent communications on synthetic data generated by embodied large language models placed in interactive teamwork scenarios. The language model supplies examples of what a helpful message looks like in a given situation, and the agents’ message space is shaped to sit in the same embedding space as those examples.

Figure 1
agent message space⏟learned  ⟷  natural language embedding space⏟fixed, external\ubrace{comm}{\text{agent message space}}{learned} \;\longleftrightarrow\; \ubrace{observe}{\text{natural language embedding space}}{fixed, external}
agent message space\text{agent message space}
the representation optimized for cooperative action
language embedding space\text{language embedding space}
the shared external semantic reference
Grounding aligns learned messages with an external semantic reference.

Read that as the general move rather than as an architecture. Learning Communication Protocols showed that a learned symbol’s meaning lives in the pair of policies, which is precisely why it does not transfer. Grounding gives the meaning somewhere else to live.

Three results, and they are worth stating separately because they answer different objections.

Task performance is maintained. Grounding constrains what the agents may learn, so the obvious worry is that it costs capability. The authors report that introducing language grounding does not reduce task performance.

Communication emerges faster. They also report that grounding accelerates the emergence of communication. This is the more interesting of the two: the bootstrapping problem from Learning Communication Protocols is that neither sender nor receiver can be correct first, and an external anchor gives both of them a target that does not depend on the other having already succeeded.

The protocols generalise zero-shot. The learned communication protocols exhibit zero-shot generalization in ad hoc teamwork scenarios with unseen partners and novel task states.

Implications for Cooperative Communication

Section titled “Implications for Cooperative Communication”

This is the question to leave the chapter with:

Should a communication protocol merely work, or should it also be understandable by agents that were not trained with it?

There is a real cost either way. An unconstrained protocol is free to find whatever encoding suits the task, and is private to the pair that found it. A grounded protocol gives up some of that freedom for a chance at being read.

Which you want depends on whether your agents will ever meet a stranger. That is the Adapt chapter.

Knowledge check

Why would grounding a message space in an external representation help with unseen partners?

Select one answer.

  • The paper starts from the problem this chapter reached: a learned protocol is usually not interpretable to humans, nor to agents not co-trained together, which limits ad hoc teamwork.
  • LangGround aligns the agents’ communication space with an embedding space derived from human natural language, grounded on synthetic data generated by embodied large language models in teamwork scenarios.
  • The authors report that grounding maintains task performance, accelerates the emergence of communication, and yields protocols that generalise zero-shot to unseen partners and novel task states.
  • The general move: when a learned convention’s meaning lives only in the pair that learned it, anchor it to something external.

Language Grounded Multi-agent Reinforcement Learning with Human-interpretable Communication, Huao Li, Hossein Nourkhiz Mahjoub, Behdad Chalaki, Vaishnav Tadiparthi, Kwonjoon Lee, Ehsan Moradi-Pari, Michael Lewis and Katia Sycara. Advances in Neural Information Processing Systems 37 (NeurIPS 2024). Proceedings page.