The Python Package
Every environment, baseline, training loop and plot used by the Colab labs is
published as one installable package, cooperative-marl-labs, so a notebook
holds the experiment and nothing else.
In this section you will
- Install the package, with or without the PyTorch extra
- Name which environment belongs to which lab
- Read a wireless observation without memorising array offsets
- Find the source for anything a notebook calls
- State what the wireless model does and does not simulate
Install
Section titled “Install”pip install cooperative-marl-labsThe speaker-listener protocol experiments in the Communicate lab train a small neural network, so they need one extra:
pip install "cooperative-marl-labs[learning]"Everything runs on a free Colab CPU. No GPU, no Ray, no RLlib, no Stable Baselines.
What is in it
Section titled “What is in it”| Environment | Lab | The question it asks |
|---|---|---|
SpeakerListenerEnv | Communicate | What can one symbol per episode buy? |
PartnerCoordinationEnv | Adapt | Did the ego learn to cooperate, or learn one partner? |
WirelessResourceAllocationEnv | Challenge Lab | Which pair of access points should share a channel? |
All three are PettingZoo ParallelEnvs and the test suite runs
pettingzoo.test.parallel_api_test against each of them, so a PettingZoo
release that changes the API fails there rather than in your notebook.
The Coordinate lab is not in the package. It runs entirely in the browser, so it has no Python to install.
The public API
Section titled “The public API”Three environments:
from cooperative_marl_labs.envs import ( SpeakerListenerEnv, PartnerCoordinationEnv, WirelessResourceAllocationEnv,)Scripted partners for the Adapt lab, each with its own seeded generator:
from cooperative_marl_labs.policies import ( FetchFirstPartner, # P(FETCH) = 0.95 CookFirstPartner, # 0.05 BalancedPartner, # 0.50 HeldOutPartner, # 0.25, evaluation only ReactivePartner, # takes whichever role the ego did not)Wireless agents, all tabular:
from cooperative_marl_labs.agents import ( RandomWirelessAgent, GreedyWirelessAgent, QLearningWirelessAgent, CommunicatingWirelessAgent,)Training, evaluation and plots:
from cooperative_marl_labs.training import ( train_independent_q_learning, train_vdn, train_communication_agents)from cooperative_marl_labs.evaluation import evaluate_agents, crossplay_matrixfrom cooperative_marl_labs.visualization import ( plot_protocol_heatmap, plot_crossplay_matrix, plot_partner_estimate, render_wireless_network, plot_wireless_comparison)Reading a wireless observation
Section titled “Reading a wireless observation”Each access point sees only its own situation, laid out as
[demand, quality per channel, interference per channel, previous channel, messages]Read it through the helpers rather than by index, so nothing breaks when communication is switched on and the vector gets longer:
from cooperative_marl_labs.envs import ( extract_demand, extract_channel_quality, extract_interference, extract_previous_channel,)
demand = extract_demand(observation)interference = extract_interference(observation, env) # pass env when comm is onenv.state() returns the centralized picture, for centralized training and for
evaluation. Passing it to an agent at execution time would break the
decentralized-execution constraint the whole resource is about.
Interventions
Section titled “Interventions”Every stress test in the labs goes through a method, so a notebook never reaches into environment attributes:
env.set_traffic("ap_2", demand=1.0)env.set_channel_quality(agent="ap_1", channel=1, quality=0.5)env.set_message_loss(0.3)env.set_communication(True)The speaker-listener environment has the same shape:
set_message_vocab_size for channel capacity, set_message_error for noise.
And PartnerCoordinationEnv has set_partner and set_partner_switch, the
second of which replaces the partner part-way through an episode without
warning.
What the wireless model is not
Section titled “What the wireless model is not”env.best_possible() enumerates all n_channels ** n_agents allocations and
returns the ceiling, which is what any result should be read against. Reporting
a reward without its ceiling is the most common way to make a mediocre
allocation look good.
Determinism
Section titled “Determinism”Every environment reseeds from reset(seed=...), every agent and partner owns
its own generator, and nothing in the package touches global NumPy random
state. Two runs with the same seeds give the same numbers.
One trap is worth naming because it produced a wrong result during development:
agents constructed with the same seed behave identically, so four
identically-seeded random access points pick the identical channel every step
and score like the worst possible policy rather than like chance.
make_fixed_agents offsets the seed per agent for exactly that reason.
Source and development
Section titled “Source and development”git clone https://github.com/rexsimiloluwah/marl-for-cooperative-environmentscd marl-for-cooperative-environments/cooperative-marl-labspython -m pip install -e ".[dev,learning]"pytestruff check .The suite covers the PettingZoo API, the physics, the interventions, seed reproducibility, the evaluation helpers, every public import, and every Python block in the package README.
The Python Package Summary
Section titled “The Python Package Summary”- One install,
cooperative-marl-labs, carries all three lab environments plus the baselines, training loops, evaluation helpers and plots. - The labs import from it, so a notebook contains the experiment rather than the infrastructure.
- The wireless model is deliberately simple and its assumptions are written down. Read every number as a measurement of that model.
- Read observations through the
extract_*helpers and run stress tests through theset_*methods, never by touching attributes.