Agents
All of these are tabular or tiny on purpose. A deep agent would hide the thing the labs are about: exactly what each agent conditions on, and exactly what it is credited with.
Import from cooperative_marl_labs.agents.
Classes
Section titled “Classes”from cooperative_marl_labs.agents import Agent
Agent(n_actions: int, seed: int | None = None) -> NoneMinimal interface. update is a no-op for agents that do not learn.
Every agent owns its own generator. Two agents constructed with the same seed behave identically, which is why the training helpers offset seeds per agent rather than sharing one.
Methods
act(observation=None, **kwargs) -> intupdate( self, observation, action, reward, next_observation=None, done: bool = False, ) -> None
WirelessAgent
Section titled “WirelessAgent”from cooperative_marl_labs.agents import WirelessAgent
WirelessAgent( agent_id: str, n_channels: int, n_agents: int = 4, communication: bool = False, seed: int | None = None,) -> NoneAn access point choosing one channel per step.
Subclasses read the observation through the extract_* helpers rather
than hardcoding offsets, so turning communication on cannot silently change
what a field means.
Methods
act(observation=None, **kwargs) -> intupdate( self, observation, action, reward, next_observation=None, done: bool = False, ) -> None
RandomAgent
Section titled “RandomAgent”from cooperative_marl_labs.agents import RandomAgent
RandomAgent(n_actions: int, seed: int | None = None) -> NonePicks uniformly among the actions, ignoring the observation.
Methods
act(observation=None, **kwargs) -> intupdate( self, observation, action, reward, next_observation=None, done: bool = False, ) -> None
RandomWirelessAgent
Section titled “RandomWirelessAgent”from cooperative_marl_labs.agents import RandomWirelessAgent
RandomWirelessAgent( agent_id: str, n_channels: int, n_agents: int = 4, communication: bool = False, seed: int | None = None,) -> NonePicks a channel uniformly at random.
Worth measuring rather than assuming: random spreading is a surprisingly hard baseline to beat with local rules, because it never synchronizes.
Methods
act(observation=None, **kwargs) -> intupdate( self, observation, action, reward, next_observation=None, done: bool = False, ) -> None
GreedyWirelessAgent
Section titled “GreedyWirelessAgent”from cooperative_marl_labs.agents import GreedyWirelessAgent
GreedyWirelessAgent( agent_id: str, n_channels: int, n_agents: int = 4, communication: bool = False, seed: int | None = None,) -> NonePicks the channel with the best quality minus the interference it measured.
Every decision it makes is individually sensible. Because every access point runs the identical deterministic rule on a similar observation, they tend to move to the same channel together, which is exactly why this agent is in the lab.
Methods
act(observation=None, **kwargs) -> intupdate( self, observation, action, reward, next_observation=None, done: bool = False, ) -> None
QLearningAgent
Section titled “QLearningAgent”from cooperative_marl_labs.agents import QLearningAgent
QLearningAgent( n_actions: int, key_fn: Callable[[Any], Hashable], alpha: float = 0.05, epsilon: float = 0.2, seed: int | None = None,) -> NoneQ-learning over whatever key key_fn produces from an observation.
Parameters
key_fn: Maps an observation to a hashable state. Choosing this is the modelling decision, so it is an argument rather than something hidden inside.
Methods
act(observation=None, greedy: bool = False, **kwargs) -> intupdate( self, observation, action, reward, next_observation=None, done: bool = True, ) -> None
One-step update with no bootstrap.values(observation) -> np.ndarray
The action-value row for this observation, created on first sight.
QLearningWirelessAgent
Section titled “QLearningWirelessAgent”from cooperative_marl_labs.agents import QLearningWirelessAgent
QLearningWirelessAgent( agent_id: str, n_channels: int, n_agents: int = 4, communication: bool = False, alpha: float = 0.05, epsilon: float = 0.2, seed: int | None = None,) -> NoneTabular Q-learning over the discretized local observation.
Methods
act(observation=None, greedy: bool = False, **kwargs) -> intupdate( self, observation, action, reward, next_observation=None, done: bool = True, ) -> Nonevalues(observation) -> np.ndarray
CommunicatingWirelessAgent
Section titled “CommunicatingWirelessAgent”from cooperative_marl_labs.agents import CommunicatingWirelessAgent
CommunicatingWirelessAgent( agent_id: str, n_channels: int, n_agents: int = 4, **kwargs,) -> NoneThe same learner, with the received message bits in its state key.
Nothing else changes, which is the point: the only difference between this agent and the one above is what it is allowed to condition on.
Methods
act(observation=None, greedy: bool = False, **kwargs) -> intupdate( self, observation, action, reward, next_observation=None, done: bool = True, ) -> Nonevalues(observation) -> np.ndarray
Speaker
Section titled “Speaker”from cooperative_marl_labs.agents import Speaker
Speaker(n_targets: int, n_messages: int, hidden: int = 32) -> NoneBase class for all neural network modules.
Your models should also subclass this class.
Modules can also contain other Modules, allowing them to be nested in a tree structure. You can assign the submodules as regular attributes::
import torch.nn as nn import torch.nn.functional as F
class Model(nn.Module): def init(self) -> None: super().init() self.conv1 = nn.Conv2d(1, 20, 5) self.conv2 = nn.Conv2d(20, 20, 5)
def forward(self, x): x = F.relu(self.conv1(x)) return F.relu(self.conv2(x))
Submodules assigned in this way will be registered, and will also have their
parameters converted when you call :meth:to, etc.
.. note::
As per the example above, an __init__() call to the parent class
must be made before assignment on the child.
:ivar training: Boolean represents whether this module is in training or evaluation mode. :vartype training: bool
Methods
forward(x: torch.Tensor) -> torch.Tensor
Listener
Section titled “Listener”from cooperative_marl_labs.agents import Listener
Listener(n_messages: int, n_targets: int, hidden: int = 32) -> NoneBase class for all neural network modules.
Your models should also subclass this class.
Modules can also contain other Modules, allowing them to be nested in a tree structure. You can assign the submodules as regular attributes::
import torch.nn as nn import torch.nn.functional as F
class Model(nn.Module): def init(self) -> None: super().init() self.conv1 = nn.Conv2d(1, 20, 5) self.conv2 = nn.Conv2d(20, 20, 5)
def forward(self, x): x = F.relu(self.conv1(x)) return F.relu(self.conv2(x))
Submodules assigned in this way will be registered, and will also have their
parameters converted when you call :meth:to, etc.
.. note::
As per the example above, an __init__() call to the parent class
must be made before assignment on the child.
:ivar training: Boolean represents whether this module is in training or evaluation mode. :vartype training: bool
Methods
forward(x: torch.Tensor) -> torch.Tensor
Functions
Section titled “Functions”discretize_observation
Section titled “discretize_observation”from cooperative_marl_labs.agents import discretize_observation
discretize_observation( observation, n_channels: int, n_agents: int = 4, communication: bool = False,) -> HashableA small hashable state key for tabular learning.
Keeps the demand level, the previous channel, and a coarse interference level per channel. Channel quality is dropped because it is constant unless an experiment degrades a channel, and including it would multiply the table for nothing.
interference_level
Section titled “interference_level”from cooperative_marl_labs.agents import interference_level
interference_level(value: float) -> intBin a continuous interference measurement into three levels.
Interference is a summed coupling rather than a count, so it has to be binned before it can index a table. Three levels: clear, a distant neighbour, a close one.
onehot
Section titled “onehot”from cooperative_marl_labs.agents import onehot
onehot(index: int, n: int) -> torch.TensorA one-hot vector of length n with position index set.
Targets and symbols are categorical, and a one-hot input keeps the network from reading an ordering into them that the task does not have.