Environments
Every environment here is a PettingZoo ParallelEnv. reset(seed) returns (observations, infos) and step(actions) returns (observations, rewards, terminations, truncations, infos), all keyed by agent name.
Import from cooperative_marl_labs.envs.
Classes
Section titled “Classes”SpeakerListenerEnv
Section titled “SpeakerListenerEnv”from cooperative_marl_labs.envs import SpeakerListenerEnv
SpeakerListenerEnv( n_targets: int = 3, message_vocab_size: int = 3, message_error: float = 0.0,) -> NoneTwo agents, one hidden target, one symbol of bandwidth.
Methods
action_space(agent: str) -> spaces.Space
The speaker chooses a symbol; the listener chooses a target.observation_space(agent: str) -> spaces.Space
The speaker sees a one-hot target; the listener sees a one-hot symbol.render() -> str
Print and return one line: the target, the symbol sent, and the one that arrived. They differ whenmessage_errorcorrupts it.reset(seed: int | None = None, options: dict | None = None)
Draw a new hidden target. Returns(observations, infos).set_message_error(p: float) -> None
Probability that the received symbol is replaced by a random one.set_message_vocab_size(k: int) -> None
Changes channel capacity. Takes effect on the next reset.step(actions: dict[str, int])
Advance one of the episode’s two steps: the speaker sends, then the listener guesses. Returns the five ParallelEnv dictionaries.
PartnerCoordinationEnv
Section titled “PartnerCoordinationEnv”from cooperative_marl_labs.envs import PartnerCoordinationEnv
PartnerCoordinationEnv( n_steps: int = 40, partner: PartnerPolicy | None = None,) -> NoneOne learning agent, one scripted partner, n_steps per episode.
Methods
action_space(agent: str) -> spaces.Space
Two actions: FETCH and COOK.observation_space(agent: str) -> spaces.Space
A one-hot of the partner’s last role, plus a slot for “nothing yet”.render() -> str
Print and return one line naming the role the partner last took.reset(seed: int | None = None, options: dict | None = None)
Start a new episode and reset the partner. Returns(observations, infos). Callset_partnerfirst.set_partner(partner: PartnerPolicy) -> None
Sets the partner used from the next reset.set_partner_switch(step: int, partner: PartnerPolicy) -> None
Replaces the partner part-way through each episode, without warning.step(actions: dict[str, int])
Take the ego’s action, ask the partner for its own, and reward the pair 1 when the two roles differ. Returns the five dictionaries.
WirelessResourceAllocationEnv
Section titled “WirelessResourceAllocationEnv”from cooperative_marl_labs.envs import WirelessResourceAllocationEnv
WirelessResourceAllocationEnv( n_agents: int = 4, n_channels: int = 3, communication: bool = False, interference_weight: float = 0.2, communication_weight: float = 0.05, traffic: str = 'skewed', n_steps: int = 8, positions: tuple[tuple[float, float], ...] | None = None, render_mode: str | None = None,) -> Nonen_agents access points choosing among n_channels channels.
Parameters
n_agents, n_channels:
Team size and how many channels they share. With more access points
than channels somebody must share, which is the point.
communication:
When true, each access point broadcasts one bit about its own demand
and receives its neighbours’ bits in its observation.
interference_weight, communication_weight:
The two shaping terms in the team reward.
traffic:
Demand regime, one of "skewed", "uniform" or "hotspot".
Demand is drawn at reset and held for the episode.
n_steps:
Allocation decisions per episode. More than one, because the
interference field reports the previous step and would otherwise stay
empty for the whole episode.
positions:
Fixed 2D coordinates, one per access point. Defaults to the lab layout
for four access points, and to an evenly spaced line otherwise.
Methods
action_space(agent: str) -> spaces.Space
One channel choice,Discrete(n_channels).best_possible() -> float
Best team reward available for the current demands.clear_traffic() -> None
Undo everyset_trafficcall.min_interference() -> float
Interference that no allocation can avoid.observation_layout() -> dict[str, slice]
Named slices into this environment’s observation vector.observation_space(agent: str) -> spaces.Space
A Box of own demand, per-channel quality, per-channel interference, previous channel, and the message bits when communication is on.outcome(actions) -> dict[str, Any]
Every metric for one joint action, without advancing the environment.render() -> str
One line of text. For the picture, userender_wireless_network.reset(seed: int | None = None, options: dict | None = None)
Draw new traffic demands and clear the allocation. Demand is held fixed for the episode. Returns(observations, infos).reset_channel_quality() -> None
Return every channel to quality 1.0.set_channel_quality( self, agent: str, channel: int, quality: float, ) -> None
Degrade or improve one channel for one access point.set_communication(on: bool) -> None
Turn the message channel on or off between episodes.set_message_loss(p: float) -> None
Probability that a broadcast bit fails to reach a neighbour.set_traffic(agent: str, demand: float) -> None
Pin one access point’s demand. Applies from the next reset.state() -> dict[str, Any]
Centralized information, for centralized training and for evaluation.step(actions: dict[str, int])
Apply one joint channel allocation. Every access point receives the same team reward, andinfoscarries the full metric dictionary from :meth:outcome. Returns the five dictionaries.
Functions
Section titled “Functions”observation_layout
Section titled “observation_layout”from cooperative_marl_labs.envs import observation_layout
observation_layout( n_channels: int, n_agents: int, communication: bool = False,) -> dict[str, slice]Named slices into an observation vector.
Read observations through this rather than hardcoding offsets, so turning communication on cannot silently change what “interference” means.
extract_demand
Section titled “extract_demand”from cooperative_marl_labs.envs import extract_demand
extract_demand(observation, source=None) -> floatThis access point’s own traffic demand.
extract_channel_quality
Section titled “extract_channel_quality”from cooperative_marl_labs.envs import extract_channel_quality
extract_channel_quality(observation, source=None) -> np.ndarrayQuality of each channel as this access point measures it.
extract_interference
Section titled “extract_interference”from cooperative_marl_labs.envs import extract_interference
extract_interference(observation, source=None) -> np.ndarrayInterference this access point measured on each channel.
It reports the PREVIOUS step, because interference is something an access point observes rather than something it knows in advance.
extract_previous_channel
Section titled “extract_previous_channel”from cooperative_marl_labs.envs import extract_previous_channel
extract_previous_channel(observation, source=None) -> intThe channel this access point chose last step, or -1 on the first step.
achievable_rate
Section titled “achievable_rate”from cooperative_marl_labs.envs import achievable_rate
achievable_rate(quality: float, interference: float) -> floatRate in bits per symbol for one access point.
interference is the summed coupling from every other access point on
the same channel, so it is a continuous quantity rather than a count.
Constants
Section titled “Constants”| Name | Value |
|---|---|
SPEAKER | 'speaker' |
LISTENER | 'listener' |
FETCH | 0 |
COOK | 1 |
ROLE_NAMES | ('FETCH', 'COOK') |