Skip to content
MARL in Cooperative Environments
Edit this page

Environments

4 min read

Every environment here is a PettingZoo ParallelEnv. reset(seed) returns (observations, infos) and step(actions) returns (observations, rewards, terminations, truncations, infos), all keyed by agent name.

Import from cooperative_marl_labs.envs.

from cooperative_marl_labs.envs import SpeakerListenerEnv
SpeakerListenerEnv(
n_targets: int = 3,
message_vocab_size: int = 3,
message_error: float = 0.0,
) -> None

Two agents, one hidden target, one symbol of bandwidth.

Methods

  • action_space(agent: str) -> spaces.Space
    The speaker chooses a symbol; the listener chooses a target.
  • observation_space(agent: str) -> spaces.Space
    The speaker sees a one-hot target; the listener sees a one-hot symbol.
  • render() -> str
    Print and return one line: the target, the symbol sent, and the one that arrived. They differ when message_error corrupts it.
  • reset(seed: int | None = None, options: dict | None = None)
    Draw a new hidden target. Returns (observations, infos).
  • set_message_error(p: float) -> None
    Probability that the received symbol is replaced by a random one.
  • set_message_vocab_size(k: int) -> None
    Changes channel capacity. Takes effect on the next reset.
  • step(actions: dict[str, int])
    Advance one of the episode’s two steps: the speaker sends, then the listener guesses. Returns the five ParallelEnv dictionaries.
from cooperative_marl_labs.envs import PartnerCoordinationEnv
PartnerCoordinationEnv(
n_steps: int = 40,
partner: PartnerPolicy | None = None,
) -> None

One learning agent, one scripted partner, n_steps per episode.

Methods

  • action_space(agent: str) -> spaces.Space
    Two actions: FETCH and COOK.
  • observation_space(agent: str) -> spaces.Space
    A one-hot of the partner’s last role, plus a slot for “nothing yet”.
  • render() -> str
    Print and return one line naming the role the partner last took.
  • reset(seed: int | None = None, options: dict | None = None)
    Start a new episode and reset the partner. Returns (observations, infos). Call set_partner first.
  • set_partner(partner: PartnerPolicy) -> None
    Sets the partner used from the next reset.
  • set_partner_switch(step: int, partner: PartnerPolicy) -> None
    Replaces the partner part-way through each episode, without warning.
  • step(actions: dict[str, int])
    Take the ego’s action, ask the partner for its own, and reward the pair 1 when the two roles differ. Returns the five dictionaries.
from cooperative_marl_labs.envs import WirelessResourceAllocationEnv
WirelessResourceAllocationEnv(
n_agents: int = 4,
n_channels: int = 3,
communication: bool = False,
interference_weight: float = 0.2,
communication_weight: float = 0.05,
traffic: str = 'skewed',
n_steps: int = 8,
positions: tuple[tuple[float, float], ...] | None = None,
render_mode: str | None = None,
) -> None

n_agents access points choosing among n_channels channels.

Parameters

n_agents, n_channels: Team size and how many channels they share. With more access points than channels somebody must share, which is the point. communication: When true, each access point broadcasts one bit about its own demand and receives its neighbours’ bits in its observation. interference_weight, communication_weight: The two shaping terms in the team reward. traffic: Demand regime, one of "skewed", "uniform" or "hotspot". Demand is drawn at reset and held for the episode. n_steps: Allocation decisions per episode. More than one, because the interference field reports the previous step and would otherwise stay empty for the whole episode. positions: Fixed 2D coordinates, one per access point. Defaults to the lab layout for four access points, and to an evenly spaced line otherwise.

Methods

  • action_space(agent: str) -> spaces.Space
    One channel choice, Discrete(n_channels).
  • best_possible() -> float
    Best team reward available for the current demands.
  • clear_traffic() -> None
    Undo every set_traffic call.
  • min_interference() -> float
    Interference that no allocation can avoid.
  • observation_layout() -> dict[str, slice]
    Named slices into this environment’s observation vector.
  • observation_space(agent: str) -> spaces.Space
    A Box of own demand, per-channel quality, per-channel interference, previous channel, and the message bits when communication is on.
  • outcome(actions) -> dict[str, Any]
    Every metric for one joint action, without advancing the environment.
  • render() -> str
    One line of text. For the picture, use render_wireless_network.
  • reset(seed: int | None = None, options: dict | None = None)
    Draw new traffic demands and clear the allocation. Demand is held fixed for the episode. Returns (observations, infos).
  • reset_channel_quality() -> None
    Return every channel to quality 1.0.
  • set_channel_quality( self, agent: str, channel: int, quality: float, ) -> None
    Degrade or improve one channel for one access point.
  • set_communication(on: bool) -> None
    Turn the message channel on or off between episodes.
  • set_message_loss(p: float) -> None
    Probability that a broadcast bit fails to reach a neighbour.
  • set_traffic(agent: str, demand: float) -> None
    Pin one access point’s demand. Applies from the next reset.
  • state() -> dict[str, Any]
    Centralized information, for centralized training and for evaluation.
  • step(actions: dict[str, int])
    Apply one joint channel allocation. Every access point receives the same team reward, and infos carries the full metric dictionary from :meth:outcome. Returns the five dictionaries.
from cooperative_marl_labs.envs import observation_layout
observation_layout(
n_channels: int,
n_agents: int,
communication: bool = False,
) -> dict[str, slice]

Named slices into an observation vector.

Read observations through this rather than hardcoding offsets, so turning communication on cannot silently change what “interference” means.

from cooperative_marl_labs.envs import extract_demand
extract_demand(observation, source=None) -> float

This access point’s own traffic demand.

from cooperative_marl_labs.envs import extract_channel_quality
extract_channel_quality(observation, source=None) -> np.ndarray

Quality of each channel as this access point measures it.

from cooperative_marl_labs.envs import extract_interference
extract_interference(observation, source=None) -> np.ndarray

Interference this access point measured on each channel.

It reports the PREVIOUS step, because interference is something an access point observes rather than something it knows in advance.

from cooperative_marl_labs.envs import extract_previous_channel
extract_previous_channel(observation, source=None) -> int

The channel this access point chose last step, or -1 on the first step.

from cooperative_marl_labs.envs import achievable_rate
achievable_rate(quality: float, interference: float) -> float

Rate in bits per symbol for one access point.

interference is the summed coupling from every other access point on the same channel, so it is a continuous quantity rather than a count.

NameValue
SPEAKER'speaker'
LISTENER'listener'
FETCH0
COOK1
ROLE_NAMES('FETCH', 'COOK')