Skip to content
MARL in Cooperative Environments
Edit this page

Partner Policies

1 min read

Each partner takes the FETCH role with its own probability, so no single observation identifies one. Telling them apart needs a history, which is what makes a partner model worth building.

Import from cooperative_marl_labs.policies.

from cooperative_marl_labs.policies import PartnerPolicy
PartnerPolicy(seed: int | None = None) -> None

Base class for a scripted partner.

Owns a seeded generator, so a partner’s behaviour is reproducible and no partner touches global NumPy random state.

Methods

  • act(observation: Any = None, history: Any = None) -> int
  • reset() -> None
    Called at every env.reset. Stateless partners need nothing.
from cooperative_marl_labs.policies import FetchFirstPartner
FetchFirstPartner(seed: int | None = None) -> None

Almost always fetches. The ego should cook.

Methods

  • act(observation: Any = None, history: Any = None) -> int
  • reset() -> None
    Called at every env.reset. Stateless partners need nothing.
from cooperative_marl_labs.policies import CookFirstPartner
CookFirstPartner(seed: int | None = None) -> None

Almost always cooks. The ego should fetch.

Methods

  • act(observation: Any = None, history: Any = None) -> int
  • reset() -> None
    Called at every env.reset. Stateless partners need nothing.
from cooperative_marl_labs.policies import BalancedPartner
BalancedPartner(seed: int | None = None) -> None

Takes each role half the time, so nothing the ego does beats 0.5.

Methods

  • act(observation: Any = None, history: Any = None) -> int
  • reset() -> None
    Called at every env.reset. Stateless partners need nothing.
from cooperative_marl_labs.policies import ReactivePartner
ReactivePartner(seed: int | None = None) -> None

Takes whichever role the ego did not take last step.

Unlike the others this one is not stochastic, so an ego that settles into a fixed role is complemented perfectly. It is in the training population to stop a learner assuming every partner is a coin flip.

Methods

  • act(observation: Any = None, history: Any = None) -> int
  • reset() -> None
    Called at every env.reset. Stateless partners need nothing.
from cooperative_marl_labs.policies import HeldOutPartner
HeldOutPartner(seed: int | None = None) -> None

Cook-leaning, and never used during training.

Methods

  • act(observation: Any = None, history: Any = None) -> int
  • reset() -> None
    Called at every env.reset. Stateless partners need nothing.
from cooperative_marl_labs.policies import all_partners
all_partners(seed: int | None = 0) -> dict[str, PartnerPolicy]

Every partner, keyed by name, each with its own offset seed.

from cooperative_marl_labs.policies import training_partners
training_partners(seed: int | None = 0) -> dict[str, PartnerPolicy]

Just the partners the ego may train against.

from cooperative_marl_labs.policies import held_out_partners
held_out_partners(seed: int | None = 0) -> dict[str, PartnerPolicy]

Just the evaluation-only partners.

NameValue
TRAINING_POPULATION('fetch-first', 'cook-first', 'balanced', 'reactive')
HELD_OUT('held-out',)