Skip to content
MARL in Cooperative Environments
Edit this page

Agents

4 min read

All of these are tabular or tiny on purpose. A deep agent would hide the thing the labs are about: exactly what each agent conditions on, and exactly what it is credited with.

Import from cooperative_marl_labs.agents.

from cooperative_marl_labs.agents import Agent
Agent(n_actions: int, seed: int | None = None) -> None

Minimal interface. update is a no-op for agents that do not learn.

Every agent owns its own generator. Two agents constructed with the same seed behave identically, which is why the training helpers offset seeds per agent rather than sharing one.

Methods

  • act(observation=None, **kwargs) -> int
  • update( self, observation, action, reward, next_observation=None, done: bool = False, ) -> None
from cooperative_marl_labs.agents import WirelessAgent
WirelessAgent(
agent_id: str,
n_channels: int,
n_agents: int = 4,
communication: bool = False,
seed: int | None = None,
) -> None

An access point choosing one channel per step.

Subclasses read the observation through the extract_* helpers rather than hardcoding offsets, so turning communication on cannot silently change what a field means.

Methods

  • act(observation=None, **kwargs) -> int
  • update( self, observation, action, reward, next_observation=None, done: bool = False, ) -> None
from cooperative_marl_labs.agents import RandomAgent
RandomAgent(n_actions: int, seed: int | None = None) -> None

Picks uniformly among the actions, ignoring the observation.

Methods

  • act(observation=None, **kwargs) -> int
  • update( self, observation, action, reward, next_observation=None, done: bool = False, ) -> None
from cooperative_marl_labs.agents import RandomWirelessAgent
RandomWirelessAgent(
agent_id: str,
n_channels: int,
n_agents: int = 4,
communication: bool = False,
seed: int | None = None,
) -> None

Picks a channel uniformly at random.

Worth measuring rather than assuming: random spreading is a surprisingly hard baseline to beat with local rules, because it never synchronizes.

Methods

  • act(observation=None, **kwargs) -> int
  • update( self, observation, action, reward, next_observation=None, done: bool = False, ) -> None
from cooperative_marl_labs.agents import GreedyWirelessAgent
GreedyWirelessAgent(
agent_id: str,
n_channels: int,
n_agents: int = 4,
communication: bool = False,
seed: int | None = None,
) -> None

Picks the channel with the best quality minus the interference it measured.

Every decision it makes is individually sensible. Because every access point runs the identical deterministic rule on a similar observation, they tend to move to the same channel together, which is exactly why this agent is in the lab.

Methods

  • act(observation=None, **kwargs) -> int
  • update( self, observation, action, reward, next_observation=None, done: bool = False, ) -> None
from cooperative_marl_labs.agents import QLearningAgent
QLearningAgent(
n_actions: int,
key_fn: Callable[[Any], Hashable],
alpha: float = 0.05,
epsilon: float = 0.2,
seed: int | None = None,
) -> None

Q-learning over whatever key key_fn produces from an observation.

Parameters

key_fn: Maps an observation to a hashable state. Choosing this is the modelling decision, so it is an argument rather than something hidden inside.

Methods

  • act(observation=None, greedy: bool = False, **kwargs) -> int
  • update( self, observation, action, reward, next_observation=None, done: bool = True, ) -> None
    One-step update with no bootstrap.
  • values(observation) -> np.ndarray
    The action-value row for this observation, created on first sight.
from cooperative_marl_labs.agents import QLearningWirelessAgent
QLearningWirelessAgent(
agent_id: str,
n_channels: int,
n_agents: int = 4,
communication: bool = False,
alpha: float = 0.05,
epsilon: float = 0.2,
seed: int | None = None,
) -> None

Tabular Q-learning over the discretized local observation.

Methods

  • act(observation=None, greedy: bool = False, **kwargs) -> int
  • update( self, observation, action, reward, next_observation=None, done: bool = True, ) -> None
  • values(observation) -> np.ndarray
from cooperative_marl_labs.agents import CommunicatingWirelessAgent
CommunicatingWirelessAgent(
agent_id: str,
n_channels: int,
n_agents: int = 4,
**kwargs,
) -> None

The same learner, with the received message bits in its state key.

Nothing else changes, which is the point: the only difference between this agent and the one above is what it is allowed to condition on.

Methods

  • act(observation=None, greedy: bool = False, **kwargs) -> int
  • update( self, observation, action, reward, next_observation=None, done: bool = True, ) -> None
  • values(observation) -> np.ndarray
from cooperative_marl_labs.agents import Speaker
Speaker(n_targets: int, n_messages: int, hidden: int = 32) -> None

Base class for all neural network modules.

Your models should also subclass this class.

Modules can also contain other Modules, allowing them to be nested in a tree structure. You can assign the submodules as regular attributes::

import torch.nn as nn import torch.nn.functional as F

class Model(nn.Module): def init(self) -> None: super().init() self.conv1 = nn.Conv2d(1, 20, 5) self.conv2 = nn.Conv2d(20, 20, 5)

def forward(self, x): x = F.relu(self.conv1(x)) return F.relu(self.conv2(x))

Submodules assigned in this way will be registered, and will also have their parameters converted when you call :meth:to, etc.

.. note:: As per the example above, an __init__() call to the parent class must be made before assignment on the child.

:ivar training: Boolean represents whether this module is in training or evaluation mode. :vartype training: bool

Methods

  • forward(x: torch.Tensor) -> torch.Tensor
from cooperative_marl_labs.agents import Listener
Listener(n_messages: int, n_targets: int, hidden: int = 32) -> None

Base class for all neural network modules.

Your models should also subclass this class.

Modules can also contain other Modules, allowing them to be nested in a tree structure. You can assign the submodules as regular attributes::

import torch.nn as nn import torch.nn.functional as F

class Model(nn.Module): def init(self) -> None: super().init() self.conv1 = nn.Conv2d(1, 20, 5) self.conv2 = nn.Conv2d(20, 20, 5)

def forward(self, x): x = F.relu(self.conv1(x)) return F.relu(self.conv2(x))

Submodules assigned in this way will be registered, and will also have their parameters converted when you call :meth:to, etc.

.. note:: As per the example above, an __init__() call to the parent class must be made before assignment on the child.

:ivar training: Boolean represents whether this module is in training or evaluation mode. :vartype training: bool

Methods

  • forward(x: torch.Tensor) -> torch.Tensor
from cooperative_marl_labs.agents import discretize_observation
discretize_observation(
observation,
n_channels: int,
n_agents: int = 4,
communication: bool = False,
) -> Hashable

A small hashable state key for tabular learning.

Keeps the demand level, the previous channel, and a coarse interference level per channel. Channel quality is dropped because it is constant unless an experiment degrades a channel, and including it would multiply the table for nothing.

from cooperative_marl_labs.agents import interference_level
interference_level(value: float) -> int

Bin a continuous interference measurement into three levels.

Interference is a summed coupling rather than a count, so it has to be binned before it can index a table. Three levels: clear, a distant neighbour, a close one.

from cooperative_marl_labs.agents import onehot
onehot(index: int, n: int) -> torch.Tensor

A one-hot vector of length n with position index set.

Targets and symbols are categorical, and a one-hot input keeps the network from reading an ordering into them that the task does not have.