1.3Independent Learning
Independent learning gives every agent its own single-agent learning process, using local observations and actions without an explicit model of the other agents.
In this section you will
- Define the independent policies each agent learns
- Explain why a learning partner appears as changing environment dynamics
- Identify the costs of non-stationarity and a weak credit signal
- State the conditions under which independent learning is a strong baseline
Independent Policies
Section titled “Independent Policies”Independent learning means each agent runs a standard single-agent algorithm on its own stream of experience. Agent sees its own observations, takes its own actions, receives the team reward, and updates its own function, as if the other agents were not there.
- agent i’s own value function, over its own actions only
- or its own policy, if the method is policy-based
- its own observation and its own action, no partner appears anywhere
Look at what is absent. No for any other agent, no joint action, no model of a partner, no shared parameters unless you choose to share them. That absence is the definition: an independent learner does not explicitly model the other agents at all.
Other Agents as Environment Dynamics
Section titled “Other Agents as Environment Dynamics”Now the central problem, and it follows directly from that absence.
If an agent does not model its partners, its partners do not disappear. They become part of what it calls “the environment”, and unlike a stove, they learn.
The textbook states this precisely: under independent learning, the other agents are effectively treated as a non-stationary part of the environment, and changes in their policies alter the transition, observation and reward functions as perceived by another agent.
Read that list again, because it is worse than it first sounds. It is not just that outcomes get noisier. All three of the things an agent is trying to learn about shift underneath it.
| From agent 1’s point of view | Because |
|---|---|
| the transition rule changed | the same action now leads somewhere else, since agent 2 contributes to every transition |
| the observations changed | agent 2 is in the kitchen, so where it stands and what it carries is part of what agent 1 sees |
| the reward rule changed | the reward is a function of the joint action, and agent 2 supplies half of it |
Concretely:
| Agent 2’s behaviour | What agent 1 learned | |
|---|---|---|
| Episode 100 | usually collects ingredients | prepare the stove and wait for delivery |
| Episode 500 | often prepares the stove itself | the same policy now produces two prepared stoves and no ingredient |
Agent 1 did not get worse. It did not change at all. The behaviour it had learned was a good response to a partner that no longer exists.
Limitations of Independent Learning
Section titled “Limitations of Independent Learning”Three limitations, and they are worth stating separately because they have different fixes.
No explicit use of information about the other agents. Even when a partner’s actions or observations are available during training, independent learning has no mechanism to use them. Sections 1.4 to 1.8 are largely about that gap.
Concurrent learning introduces non-stationarity. Every agent is simultaneously adapting to a target that other agents are moving. Old experience describes a team that has since changed, so replay buffers age badly and convergence guarantees from single-agent theory do not carry over.
Environmental randomness and partner change are indistinguishable. An agent has no way to attribute a surprising outcome to the world rather than to its team, so it cannot know whether to keep exploring or to re-learn.
Practical Advantages
Section titled “Practical Advantages”Now the correction, because independent learning is routinely dismissed too quickly.
It is simple. No new algorithm. If you can run PPO, you can run IPPO today.
It scales in the way that matters. Each agent learns over its own action set, so nothing here grows exponentially with the number of agents, which is exactly the wall that centralized joint-action control runs into in the next section.
It is decentralized by construction. There is no training-time infrastructure to dismantle before deployment, and no question about whether the learned policies will run locally. They already do.
Knowledge check
Correct.
Not quite.
Experience collected earlier describes a team that has since changed, so learning from it can point the agent in the wrong direction.
Right. Single-agent methods assume samples describe a fixed environment. Here the environment includes learners, so stored experience decays in validity, which is why replay buffers and old on-policy data are more dangerous in multi-agent settings.
The environment becomes random, so the agent cannot predict outcomes.
Single-agent RL handles randomness perfectly well, a stochastic transition function is standard. The problem is not randomness but drift: the distribution itself moves, so past samples describe a different problem rather than a noisy version of the same one.
Each agent’s action space grows as other agents learn.
Action spaces are fixed. What changes is the outcome of choosing from them, because partners supply the rest of the joint action.
It makes the reward function impossible to learn.
Too strong. The reward function R(s, a) is fixed; what shifts is the reward an agent sees for its own action, since the other half of the joint action keeps changing. Difficult, not impossible, independent learners often do converge to something good.
Explanation
The distinction between “noisy” and “drifting” is what separates a hard single-agent problem from a multi-agent one.
Independent Learning Summary
Section titled “Independent Learning Summary”- Independent learning runs a single-agent algorithm per agent, on local observations and local actions, with no explicit model of partners.
- Partners therefore become part of each agent’s environment, a part that learns, making the perceived transition, observation and reward functions all non-stationary.
- An agent cannot distinguish environmental randomness from a partner changing policy.
- Its costs: no use of information about other agents, non-stationarity from concurrent learning, and unattributable surprise.
- Its strengths: simplicity, scalability in the number of agents, and decentralization by construction.
- It remains a serious baseline, sometimes competitive with far more sophisticated methods. Treat it as the bar to clear, not the strawman.