Partner Policies
Each partner takes the FETCH role with its own probability, so no single observation identifies one. Telling them apart needs a history, which is what makes a partner model worth building.
Import from cooperative_marl_labs.policies.
Classes
Section titled “Classes”PartnerPolicy
Section titled “PartnerPolicy”from cooperative_marl_labs.policies import PartnerPolicy
PartnerPolicy(seed: int | None = None) -> NoneBase class for a scripted partner.
Owns a seeded generator, so a partner’s behaviour is reproducible and no partner touches global NumPy random state.
Methods
act(observation: Any = None, history: Any = None) -> intreset() -> None
Called at everyenv.reset. Stateless partners need nothing.
FetchFirstPartner
Section titled “FetchFirstPartner”from cooperative_marl_labs.policies import FetchFirstPartner
FetchFirstPartner(seed: int | None = None) -> NoneAlmost always fetches. The ego should cook.
Methods
act(observation: Any = None, history: Any = None) -> intreset() -> None
Called at everyenv.reset. Stateless partners need nothing.
CookFirstPartner
Section titled “CookFirstPartner”from cooperative_marl_labs.policies import CookFirstPartner
CookFirstPartner(seed: int | None = None) -> NoneAlmost always cooks. The ego should fetch.
Methods
act(observation: Any = None, history: Any = None) -> intreset() -> None
Called at everyenv.reset. Stateless partners need nothing.
BalancedPartner
Section titled “BalancedPartner”from cooperative_marl_labs.policies import BalancedPartner
BalancedPartner(seed: int | None = None) -> NoneTakes each role half the time, so nothing the ego does beats 0.5.
Methods
act(observation: Any = None, history: Any = None) -> intreset() -> None
Called at everyenv.reset. Stateless partners need nothing.
ReactivePartner
Section titled “ReactivePartner”from cooperative_marl_labs.policies import ReactivePartner
ReactivePartner(seed: int | None = None) -> NoneTakes whichever role the ego did not take last step.
Unlike the others this one is not stochastic, so an ego that settles into a fixed role is complemented perfectly. It is in the training population to stop a learner assuming every partner is a coin flip.
Methods
act(observation: Any = None, history: Any = None) -> intreset() -> None
Called at everyenv.reset. Stateless partners need nothing.
HeldOutPartner
Section titled “HeldOutPartner”from cooperative_marl_labs.policies import HeldOutPartner
HeldOutPartner(seed: int | None = None) -> NoneCook-leaning, and never used during training.
Methods
act(observation: Any = None, history: Any = None) -> intreset() -> None
Called at everyenv.reset. Stateless partners need nothing.
Functions
Section titled “Functions”all_partners
Section titled “all_partners”from cooperative_marl_labs.policies import all_partners
all_partners(seed: int | None = 0) -> dict[str, PartnerPolicy]Every partner, keyed by name, each with its own offset seed.
training_partners
Section titled “training_partners”from cooperative_marl_labs.policies import training_partners
training_partners(seed: int | None = 0) -> dict[str, PartnerPolicy]Just the partners the ego may train against.
held_out_partners
Section titled “held_out_partners”from cooperative_marl_labs.policies import held_out_partners
held_out_partners(seed: int | None = 0) -> dict[str, PartnerPolicy]Just the evaluation-only partners.
Constants
Section titled “Constants”| Name | Value |
|---|---|
TRAINING_POPULATION | ('fetch-first', 'cook-first', 'balanced', 'reactive') |
HELD_OUT | ('held-out',) |