Skip to content
MARL in Cooperative Environments
Edit this page

3.12Adapt Lab

7 min read

The Adapt Lab is a hands-on Colab investigation of cooperation with familiar and unseen partners in a small kitchen. You will compare a specialist, a generalist trained with diverse partners, and an adaptive agent that identifies behaviour online. The experiments examine familiar performance, two different held-out partners, generalization gaps, mid-episode partner change, and cross-play. Supplied tabular policies and short episodes run on free Colab CPU, so the work centers on interpreting partner-dependent behaviour.

02_adapt.ipynb
Time
about 30 minutes
Compute
CPU only, nothing to install

Two agents. Each step, each picks FETCH, COOK or WAIT.

  • Exactly one FETCH obtains an ingredient.
  • Exactly one COOK on an ingredient serves an order: +1.
  • Both FETCH: they collide at the store, nothing obtained.
  • Both COOK: they both grab the pan, the dish is ruined.

The collisions are what make it a coordination problem. The pair must split the roles and nothing says which agent takes which, so the split is an arbitrary convention. An episode is 40 steps, so a specialised pair serves about 39 orders.

PartnerBehaviourIn training?
ingredient-firstalways fetchesyes
cooking-firstalways cooksyes
reactivewaits, then does whatever you did notyes
idle-then-cookcooks if an ingredient is there, else waitsyes
mostly-fetchfetches 85% of the time, else waitsheld out
alternatealternates fetch and cookheld out

The two held-out partners differ from each other in an important way. mostly-fetch behaves much like ingredient-first, so it is a stranger near the training distribution. alternate needs a response no training partner needed, so it is a stranger outside it.

That is the ZSC-Eval point from Partner Dependence0 built into the experiment: how much a held-out partner resembles the training set decides what holding it out measures.

Specialist, trained against ingredient-first alone. Standard practice: fix a partner, optimise team return.

Generalist, trained against all four population partners.

Adaptive, which watches for four steps, works out which known type the partner most resembles, then plays that type’s best response. Agent Modelling’s agent modelling in its simplest form, with no weight updates at deployment: it revises a belief and switches which fixed policy it follows.

The ceiling column is a separate agent trained against that partner alone, so it is what any method could achieve knowing exactly who it faced.

PartnerCeilingSpecialistGeneralistAdaptive
ingredient-first39393836
cooking-first39383935
reactive3903937
idle-then-cook39383935
mostly-fetch (held out)33292825
alternate (held out)39002

Four things to take from that table.

The specialist scores zero against reactive. Not badly, zero. It learned “my partner fetches, so I cook”, which is optimal against ingredient-first and catastrophic against a partner that is also waiting for someone else to commit. Nothing is broken in either agent.

Diversity removes the failure. The generalist reaches 38 or 39 across the whole population. It works exactly as Training Partner Diversity described, by removing a shortcut: “my partner fetches” stops paying once half the population does not fetch. It costs a point against ingredient-first, which is the price of not exploiting one partner’s habits.

The specialist beats the generalist on mostly-fetch. 29 against 28. The stranger happens to resemble its training partner, so its narrow rule transfers. A specialist is not always worse on unseen partners, only on unseen partners unlike its own.

Nobody handles alternate. All three sit near zero against a ceiling of 39. No method here invents a response that no training partner required.

The adaptive agent’s belief is the interesting part. The notebook prints which partner it thinks it has, after each step of observation:

True partnerWhat the agent concludes
ingredient-firstingredient-first, correct
cooking-firstcooking-first, correct
reactivereactive, correct
mostly-fetchingredient-first: wrong, and useful
alternateingredient-first: wrong, and useless

The two held-out rows are the lesson. For mostly-fetch, the agent picks the nearest thing it knows and the response is nearly right. For alternate, it picks the nearest thing it knows and the response is simply wrong, because nothing it knows behaves like that.

Now the result worth the whole lab. Compute the Partner Generalization diagnostic, familiar minus unseen, for each arm.

ArmFamiliarUnseenGap
Specialist28.814.314.4
Generalist38.814.124.6
Adaptive35.813.722.1

The specialist has the smallest gap, and it is the worst agent here.

Its gap is small because its familiar-partner average is dragged down by the zero against reactive, not because it generalises. This is the warning from Evaluating Partner Generalization, measured: uniform mediocrity has a small gap.

The notebook’s last experiment swaps the partner at step 20 without telling the agent. Switching from ingredient-first to cooking-first:

ArmWith a switchNo switchCost
Specialist37392
Generalist36382
Adaptive34362

Read this carefully, because the honest reading is not dramatic. The switch costs every arm about two orders, so none of them collapses: all three condition on what the partner did last step, and both partners here are handled by reacting.

The adaptive agent is nonetheless the weakest on the switch, which is worth sitting with given that adaptation is its whole purpose. The reason is in its code: it commits once, after four steps, and never revises. A belief formed early and held forever is not adaptation, it is a slower form of assumption.

The notebook asks you to make it re-identify continuously, and to find what that costs on partners that never change.

For comparison, here is the Communicate chapter cross-play matrix again, now with the vocabulary to read it properly. Switch to gap mode and note that every agent looks identical on that statistic, because they are all perfect with themselves and all poor with each other.

R0R1R2R3R4R5
S0
S1
S2
S3
S4
S5

Knowledge check

Selecting a method by minimising the generalization gap would choose the specialist in this lab. What does that tell you about the gap?

Select one answer.

Run it yourself:Open in Colab

  • A specialist trained with one partner scored zero with a perfectly competent one it never met.
  • Diversity removed the failure and cost about one point of peak performance per familiar partner.
  • Modelling works when the stranger resembles someone known. It fails confidently, not gracefully, when it does not.
  • Held-out partners near the training distribution and outside it measure very different things. mostly-fetch and alternate differ by 25 orders for the adaptive agent.
  • The generalization gap is not a target. The worst agent here has the smallest one.
  • An agent that commits to a belief once and never revises is not adapting.