Skip to content
MARL in Cooperative Environments
Edit this page

3.7Ad Hoc Teamwork

5 min read

Ad hoc teamwork studies cooperation with partners that were not selected, designed, or co-trained with the controlled agent.

In this section you will

  • Define the unfamiliar-partner setting
  • Contrast a fixed team with an ad hoc team
  • Describe the open deployment conditions this reflects
  • State the three requirements this places on a policy

Everything in Chapters 1 and 2 quietly assumed the opposite. The agents were trained together, by us, under one objective, and deployed together. Under that assumption a convention is a perfectly good solution, because both halves of it were installed by the same process.

Standard cooperative MARLAd hoc teamwork
TrainingA with B, every episodeA with B, then C, then D
DeploymentA with BA with ?
Who designed the partnerswe didsomebody else, or nobody we can ask
Can we inspect their policiesyesno
Do they share our objectiveby constructionwe hope so
Is a convention a solutionyesonly if the other party holds it too

The last row is the one that reorganises everything. A convention is a solution exactly when you control both parties, and ad hoc teamwork is defined by not controlling both.

The key idea to carry, whatever the formulation:

The learner cannot assume that it designed, trained, or knows the policies of everyone else.

Not a thought experiment. The situations where cooperative multi-agent systems are most wanted are mostly ad hoc:

Autonomous vehicles from different manufacturers. Each fleet is trained separately, by companies that do not share models, and they meet at junctions. No manufacturer can train against the others’ actual policies.

Robots from different organisations. Warehouse equipment, agricultural machinery and inspection drones increasingly come from multiple vendors and have to share space.

AI agents built by different developers. Two assistants negotiating a meeting were not co-trained and cannot be.

Human-AI teams. The most common case, and the hardest, because the human was not trained at all in the relevant sense and will not adapt to an arbitrary convention just because the system holds one.

Emergency response. Teams assembled from whatever units are available, with no opportunity to train the combination beforehand.

Three requirements follow, and each one has already appeared:

Do not depend on an arbitrary convention. If your behaviour requires the partner to hold an agreement, and the agreement was never communicated, you are relying on coincidence. The Communication Lab measured the coincidence rate.

Infer what you can from behaviour. A stranger’s policy is unavailable and its behaviour is observable. Sections 3.4 and 3.5 are the machinery.

Be evaluated against strangers. An ad hoc teamwork claim tested with familiar partners is not a claim about ad hoc teamwork. Evaluating Partner Generalization makes this a procedure.

Knowledge check

A team of four warehouse robots is trained together and deployed together, in the same warehouse, forever. Is this ad hoc teamwork?

Select one answer.

  • Ad hoc teamwork is collaboration with previously unknown agents whose behaviours are initially unknown: cooperate without having trained together.
  • Chapters 1 and 2 assumed co-training, which is what makes an arbitrary convention a legitimate solution.
  • Many formulations assume the learner controls one agent, with the rest external. Real deployments often control some and not others.
  • The learner cannot assume it designed, trained, or knows the policies of everyone else.
  • The realistic cases are mostly ad hoc: vehicles across manufacturers, robots across vendors, agents across developers, human-AI teams, emergency response.
  • The shared constraint is interoperability, and it is not addressed by more training against a familiar partner.