3.10Research Connection: N-Agent Ad Hoc Teamwork
N-Agent Ad Hoc Teamwork extends partner generalization to teams in which both the number and behavioural types of uncontrolled agents can change.
In this section you will
- Locate N-Agent Ad Hoc Teamwork between controlling one agent and the whole team
- Explain why variable team composition matters at deployment
- Describe POAM’s partner-representation approach
- Read the reported findings and what they do and do not establish
The N-Agent Ad Hoc Teamwork Setting
Section titled “The N-Agent Ad Hoc Teamwork Setting”Two settings have been well studied, and they are the two extremes.
| How many agents the learner controls | |
|---|---|
| Standard cooperative MARL | all of them |
| Traditional ad hoc teamwork | one of them |
N-Agent Ad Hoc Teamwork (NeurIPS 2024) points out that real systems can sit between these extremes. The authors introduce a setting in which a set of autonomous agents must interact and cooperate with dynamically varying numbers and types of partners.
Both halves of that phrase matter, and the first is the one previous framings did not cover.
- the policies available to the learner
- partners whose number and types may vary
Variable Team Composition
Section titled “Variable Team Composition”Return to the motivating examples from Ad Hoc Teamwork and count.
A warehouse operator deploying twelve of its own robots alongside a subcontractor’s eight controls some of the team. So does a fleet operator whose vehicles meet other manufacturers’ vehicles at a junction: how many of each is at the junction changes minute to minute.
Neither situation is described by controlling one agent, and neither is described by controlling all of them. And the difference is not cosmetic: an agent that can rely on three co-trained colleagues behaves differently from one that is alone among strangers, and in the middle case it does not know in advance which it will be.
POAM and Partner Representations
Section titled “POAM and Partner Representations”The authors propose Policy Optimization with Agent Modelling, or POAM.
The idea connects directly to sections 3.4 and 3.5 rather than arriving as an unrelated algorithm. POAM is a policy gradient method that learns representations of partner behaviours and uses them to adapt.
That is the behaviour embedding from Partner Representations, put to work:
- locally available evidence about the partners
- a learned summary of the variable partner set
- the controlled agent’s response
Note what the representation has to survive here. Under NAHT the partners vary in kind and in number, so a representation indexed by partner identity is doubly useless: it has no entry for a stranger, and no way to describe a team of three strangers rather than one.
Reported Experimental Findings
Section titled “Reported Experimental Findings”The authors evaluate POAM in multi-agent particle environments and StarCraft II, and report improved cooperative performance over baselines along with generalization to previously unseen partners.
Implications for Ad Hoc Teamwork
Section titled “Implications for Ad Hoc Teamwork”How should an agent represent the behaviour of partners when both their type and their number can change?
Worth sitting with. A representation of one partner is a vector. A representation of a team of unknown size, whose members you may be individually uncertain about, is a harder object, and how to build one is open.
Knowledge check
Correct.
Not quite.
Because the partners vary in both kind and number, so an identity index has no entry for a stranger and no way to describe a team of several unknowns.
Both failures at once. *Partner Representations* established that identity representations cannot place a stranger. NAHT adds that the team composition itself varies, so even a behaviour representation has to summarise a set of partners of unknown size.
Because identities are private information that agents cannot observe.
Identity is sometimes observable, and that is not the problem. The problem is that knowing which partner you have only helps if you have met it before.
Because policy gradient methods cannot use discrete representations.
They can. The obstacle is what the representation can express about unseen partners, not how it is optimised.
Because NAHT requires controlling only one agent.
The opposite: NAHT is specifically about controlling *some* of the team, between the one-agent and all-agents extremes.
Explanation
The pattern to carry: whenever the thing you must generalise over can vary in size as well as in kind, a lookup table is even further from sufficient.
N-Agent Ad Hoc Teamwork Lessons
Section titled “N-Agent Ad Hoc Teamwork Lessons”- Two well-studied extremes: the learner controls all agents (standard cooperative MARL) or one (traditional ad hoc teamwork).
- NAHT is the setting in between, where agents cooperate with dynamically varying numbers and types of partners.
- Real deployments are mostly in the middle: your own fleet alongside somebody else’s, in proportions that change.
- POAM is a policy gradient method that learns representations of partner behaviour and uses them to adapt, which is the Partner Representations mechanism instantiated.
- The authors report improved cooperative performance and generalization to unseen partners on multi-agent particle environments and StarCraft II.
Further reading
Section titled “Further reading”N-Agent Ad Hoc Teamwork, Caroline Wang, Arrasy Rahman, Ishan Durugkar, Elad Liebman and Peter Stone. Advances in Neural Information Processing Systems 37 (NeurIPS 2024). Proceedings page.