3.13Adapt Worksheet
This five-part digital activity consolidates the chapter’s approach to partner generalization. You will match adaptation concepts to their applications, calculate a familiar-to-unseen generalization gap, and analyze a cross-play table. An interactive graph compares familiar and held-out partner performance for specialist, generalist, and adaptive agents. The final prompt asks you to weigh specialization against robust deployment performance. Every earlier item is automatically checkable.
What you will be able to do
- Match partner-generalization concepts to their use in cooperative MARL. Understand
- Calculate simple measures of familiar and unseen-partner performance. Apply
- Interpret results across different partners. Analyse
- Evaluate the trade-off between specialization and partner generalization. Evaluate
1 · Match Concepts to Applications
Section titled “1 · Match Concepts to Applications”Drag each block into a slot, or tap one and then tap a slot.
Train with several different partner policies.
Infer how another agent is likely to behave.
Cooperate with previously unknown agents.
Evaluate with an unseen partner, without prior joint training.
Pair policies that were not trained together and measure the result.
2 · Calculate the Generalization Gap
Section titled “2 · Calculate the Generalization Gap”The worksheet diagnostic from Partner Generalization: .
Agent A scores and . Calculate .
Agent B scores and . Calculate its gap.
3 · Calculate Cross-Play Performance
Section titled “3 · Calculate Cross-Play Performance”Percentage success. The diagonal is each policy with its own training partner.
| P1 | P2 | P3 | P4 | |
|---|---|---|---|---|
| A | 94 | 81 | 65 | 72 |
| B | 80 | 93 | 84 | 79 |
| C | 61 | 86 | 95 | 88 |
| D | 70 | 82 | 85 | 96 |
Calculate policy B’s mean over the three partners it did not train with.
Exclude the diagonal entry. .
Now the same for policy A.
4 · Explore Familiar and Unseen-Partner Performance
Section titled “4 · Explore Familiar and Unseen-Partner Performance”Real results from the Adapt Lab: three agents, 40-step episodes, orders
served. mostly-fetch is a held-out partner that behaves much like a training
partner. alternate is a held-out partner that does not.
- Familiar mean
- mostly-fetch (near)
- alternate (far)
By how much does the Generalist beat the Specialist on familiar partners?
5 · Evaluate the Specialization Trade-off
Section titled “5 · Evaluate the Specialization Trade-off”A fixed-partner agent achieves 96% with its training partner and 54% with unseen partners.
A diverse-partner agent achieves 90% and 82% respectively.
Which would you deploy in a system where partners change frequently, and what trade-off are you accepting?
AnswersReveal
Partner diversity · agent modelling · ad hoc teamwork · zero-shot coordination · cross-play, matched to the descriptions in order.
.
.
.
.
orders.
No single right answer.
The diverse agent, in a system where partners change frequently. 82 against 54 on unseen partners is a 28-point difference, and unseen partners are the common case by assumption.
The trade-off: 6 points of peak performance with a familiar partner, 96 down to 90. You are giving up the ability to exploit one partner’s habits in exchange for competence across a range, which is the cost Training Partner Diversity named.
A strong answer also notes what the numbers omit: how the unseen partners were chosen. If they resemble the training partners, both figures are optimistic and the 82 in particular cannot be trusted.
Review with Flashcards
Section titled “Review with Flashcards”Revisit the 10 key concepts, equations, and intuitions from this chapter.