Skip to content
MARL in Cooperative Environments
Edit this page

3.10Research Connection: N-Agent Ad Hoc Teamwork

5 min read

N-Agent Ad Hoc Teamwork extends partner generalization to teams in which both the number and behavioural types of uncontrolled agents can change.

In this section you will

  • Locate N-Agent Ad Hoc Teamwork between controlling one agent and the whole team
  • Explain why variable team composition matters at deployment
  • Describe POAM’s partner-representation approach
  • Read the reported findings and what they do and do not establish

Two settings have been well studied, and they are the two extremes.

How many agents the learner controls
Standard cooperative MARLall of them
Traditional ad hoc teamworkone of them

N-Agent Ad Hoc Teamwork (NeurIPS 2024) points out that real systems can sit between these extremes. The authors introduce a setting in which a set of autonomous agents must interact and cooperate with dynamically varying numbers and types of partners.

Both halves of that phrase matter, and the first is the one previous framings did not cover.

Figure 1
agents we control⏟some  +  agents we do not⏟how many? which kinds?\underbrace{\tone{policy}{\text{agents we control}}}_{\text{some}} \;+\; \underbrace{\tone{conflict}{\text{agents we do not}}}_{\text{how many? which kinds?}}
controlled agents\text{controlled agents}
the policies available to the learner
uncontrolled agents\text{uncontrolled agents}
partners whose number and types may vary
NAHT teams combine controlled policies with a variable set of unfamiliar partners.

Return to the motivating examples from Ad Hoc Teamwork and count.

A warehouse operator deploying twelve of its own robots alongside a subcontractor’s eight controls some of the team. So does a fleet operator whose vehicles meet other manufacturers’ vehicles at a junction: how many of each is at the junction changes minute to minute.

Neither situation is described by controlling one agent, and neither is described by controlling all of them. And the difference is not cosmetic: an agent that can rely on three co-trained colleagues behaves differently from one that is alone among strangers, and in the middle case it does not know in advance which it will be.

The authors propose Policy Optimization with Agent Modelling, or POAM.

The idea connects directly to sections 3.4 and 3.5 rather than arriving as an unrelated algorithm. POAM is a policy gradient method that learns representations of partner behaviours and uses them to adapt.

That is the behaviour embedding from Partner Representations, put to work:

Figure 2
observed partner behaviour⏟what I have seen→learned representation⏟who these partners are→my action⏟what I should do about it\ubrace{observe}{\text{observed partner behaviour}}{what I have seen} \rightarrow \ubrace{policy}{\text{learned representation}}{who these partners are} \rightarrow \ubrace{action}{\text{my action}}{what I should do about it}
observed behaviour\text{observed behaviour}
locally available evidence about the partners
representation\text{representation}
a learned summary of the variable partner set
action\text{action}
the controlled agent’s response
POAM conditions controlled actions on learned representations of the observed partner set.

Note what the representation has to survive here. Under NAHT the partners vary in kind and in number, so a representation indexed by partner identity is doubly useless: it has no entry for a stranger, and no way to describe a team of three strangers rather than one.

The authors evaluate POAM in multi-agent particle environments and StarCraft II, and report improved cooperative performance over baselines along with generalization to previously unseen partners.

How should an agent represent the behaviour of partners when both their type and their number can change?

Worth sitting with. A representation of one partner is a vector. A representation of a team of unknown size, whose members you may be individually uncertain about, is a harder object, and how to build one is open.

Knowledge check

Why is a partner representation indexed by identity especially useless in the NAHT setting?

Select one answer.

  • Two well-studied extremes: the learner controls all agents (standard cooperative MARL) or one (traditional ad hoc teamwork).
  • NAHT is the setting in between, where agents cooperate with dynamically varying numbers and types of partners.
  • Real deployments are mostly in the middle: your own fleet alongside somebody else’s, in proportions that change.
  • POAM is a policy gradient method that learns representations of partner behaviour and uses them to adapt, which is the Partner Representations mechanism instantiated.
  • The authors report improved cooperative performance and generalization to unseen partners on multi-agent particle environments and StarCraft II.

N-Agent Ad Hoc Teamwork, Caroline Wang, Arrasy Rahman, Ishan Durugkar, Elad Liebman and Peter Stone. Advances in Neural Information Processing Systems 37 (NeurIPS 2024). Proceedings page.