Training
One loop covers independent learning and value decomposition, because the only difference between them is how the error is computed. Keeping them in one function is what makes the comparison honest.
Import from cooperative_marl_labs.training.
Functions
Section titled “Functions”train_independent_q_learning
Section titled “train_independent_q_learning”from cooperative_marl_labs.training import train_independent_q_learning
train_independent_q_learning( env: WirelessResourceAllocationEnv, episodes: int = 4000, seed: int = 0, epsilon: float = 0.2, alpha: float = 0.05,) -> tuple[dict[str, QLearningWirelessAgent], list[float]]Independent Q-learning: every access point learns alone.
Each one receives the whole team reward as its own target, so it cannot tell its own contribution apart from its neighbours’.
Returns the trained agents and the per-episode mean team reward.
train_vdn
Section titled “train_vdn”from cooperative_marl_labs.training import train_vdn
train_vdn( env: WirelessResourceAllocationEnv, episodes: int = 4000, seed: int = 0, epsilon: float = 0.2, alpha: float = 0.05,) -> tuple[dict[str, QLearningWirelessAgent], list[float]]Value decomposition: one error on the sum of the individual values.
Execution stays decentralized. Only the target changes, which is the whole of VDN at this scale.
Returns the trained agents and the per-episode mean team reward.
combine_values
Section titled “combine_values”from cooperative_marl_labs.training import combine_values
combine_values(q_values: Sequence[float]) -> floatQ_tot as the sum of the per-agent values. This is the VDN assumption.
make_wireless_agents
Section titled “make_wireless_agents”from cooperative_marl_labs.training import make_wireless_agents
make_wireless_agents( env: WirelessResourceAllocationEnv, epsilon: float = 0.2, alpha: float = 0.05, seed: int = 0,) -> dict[str, QLearningWirelessAgent]One learner per access point, matching the environment’s comm setting.
make_fixed_agents
Section titled “make_fixed_agents”from cooperative_marl_labs.training import make_fixed_agents
make_fixed_agents( env: WirelessResourceAllocationEnv, cls, seed: int = 0, **kwargs,) -> dict[str, Any]One non-learning agent per access point, each with a DIFFERENT seed.
Sharing one seed across agents is a trap worth avoiding: identically seeded random agents draw the identical channel every step, so they collide on every step and score like the worst possible policy rather than like chance.
train_communication_agents
Section titled “train_communication_agents”from cooperative_marl_labs.training import train_communication_agents
train_communication_agents( n_targets: int = 3, n_messages: int = 3, message_error: float = 0.0, episodes: int = 4000, batch: int = 32, lr: float = 0.01, seed: int = 0, hidden: int = 32,) -> tuple[Speaker, Listener, list[float]]Train a speaker and a listener from the shared task reward alone.
Neither agent is told what a symbol should mean. Whatever convention comes out is theirs, which is the point of the experiment.
Returns the trained pair and the per-batch success history.
protocol_matrix
Section titled “protocol_matrix”from cooperative_marl_labs.training import protocol_matrix
protocol_matrix(speaker: Speaker, n_targets: int) -> np.ndarrayP(message | target), one row per target.