Every card shows its question first. Reveal one at a time to test yourself, or
reveal all of them to use the deck as a revision sheet.
10 cards
Concept
What does a policy represent in reinforcement learning?
Concept
A policy defines how an agent chooses actions from the information available to it.
at∼π(at∣ot)
The policy π maps observation ot to a distribution over possible actions.
Equation
What does this mean?
at∼π(at∣ot)
Equation
At time t, the agent chooses action at according to policy π, conditioned on its current observation ot.
π
the decision rule the agent follows
ot
everything the agent can see at this step
∼
sampled from, so the policy may be stochastic
Concept
What changes when reinforcement learning moves from one agent to multiple agents?
Concept
An agent's outcome can now depend on what other agents do. Other agents may also be learning, changing the environment each agent experiences.
Concept
What is a joint action, and why is it important in MARL?
Concept
A joint action is the combination of all agents' actions at a particular time.
at=(at1,…,atn)
The environment responds to this combination, so the outcome may depend on how the agents' actions fit together.
Equation
What does this represent?
at=(at1,…,atn)
Equation
The joint action at time t, containing the actions selected by all n agents.
ati
agent i's own action, chosen from its own observation
at
the tuple the environment actually responds to
Distinction
What is the difference between the environment state and an agent's observation?
Distinction
The state represents the underlying condition of the environment. An observation contains only the information available to a particular agent.
Equation
What does this tell us about agent i's information?
oti=Oi(st)
Equation
Agent i's observation oti is generated from the underlying state st through its observation function Oi.
The agent may therefore see only part of the complete state.
Concept
What does partial observability mean in a multi-agent environment?
Concept
An agent does not directly observe all information about the environment that may be relevant to its decision.
Intuition
Why can agents in the same environment make decisions from different information?
Intuition
Each agent may have its own observation function, location, sensors, or access to information. They share the environment without necessarily sharing the same view of it.
Distinction
What does giving all agents the same reward accomplish, and what does it not guarantee?
Distinction
A shared reward aligns the agents toward the same objective.
It does not guarantee that their individual actions will be coordinated, or reveal how each agent contributed to the outcome.