Skip to content
MARL in Cooperative Environments
Edit this page

1.3Independent Learning

6 min read

Independent learning gives every agent its own single-agent learning process, using local observations and actions without an explicit model of the other agents.

In this section you will

  • Define the independent policies each agent learns
  • Explain why a learning partner appears as changing environment dynamics
  • Identify the costs of non-stationarity and a weak credit signal
  • State the conditions under which independent learning is a strong baseline

Independent learning means each agent runs a standard single-agent algorithm on its own stream of experience. Agent i\ag sees its own observations, takes its own actions, receives the team reward, and updates its own function, as if the other agents were not there.

Figure 1
Qi(oti, ati)orπi(ati∣oti)\tone{policy}{Q_\ag\bigl(\tone{observe}{\obs{\ag}},\ \tone{action}{\act{\ag}}\bigr)} \qquad\text{or}\qquad \tone{policy}{\pol{\ag}\bigl(\tone{action}{\act{\ag}} \given \tone{observe}{\obs{\ag}}\bigr)}
QiQ_\ag
agent i’s own value function, over its own actions only
πi\pol{\ag}
or its own policy, if the method is policy-based
oti,ati\obs{\ag}, \act{\ag}
its own observation and its own action, no partner appears anywhere
Every function an independent learner holds is defined on its own observation and its own action.

Look at what is absent. No atj\act{j} for any other agent, no joint action, no model of a partner, no shared parameters unless you choose to share them. That absence is the definition: an independent learner does not explicitly model the other agents at all.

Now the central problem, and it follows directly from that absence.

If an agent does not model its partners, its partners do not disappear. They become part of what it calls “the environment”, and unlike a stove, they learn.

The textbook states this precisely: under independent learning, the other agents are effectively treated as a non-stationary part of the environment, and changes in their policies alter the transition, observation and reward functions as perceived by another agent.

Read that list again, because it is worse than it first sounds. It is not just that outcomes get noisier. All three of the things an agent is trying to learn about shift underneath it.

From agent 1’s point of viewBecause
the transition rule changedthe same action now leads somewhere else, since agent 2 contributes to every transition
the observations changedagent 2 is in the kitchen, so where it stands and what it carries is part of what agent 1 sees
the reward rule changedthe reward is a function of the joint action, and agent 2 supplies half of it

Concretely:

Agent 2’s behaviourWhat agent 1 learned
Episode 100usually collects ingredientsprepare the stove and wait for delivery
Episode 500often prepares the stove itselfthe same policy now produces two prepared stoves and no ingredient

Agent 1 did not get worse. It did not change at all. The behaviour it had learned was a good response to a partner that no longer exists.

Three limitations, and they are worth stating separately because they have different fixes.

No explicit use of information about the other agents. Even when a partner’s actions or observations are available during training, independent learning has no mechanism to use them. Sections 1.4 to 1.8 are largely about that gap.

Concurrent learning introduces non-stationarity. Every agent is simultaneously adapting to a target that other agents are moving. Old experience describes a team that has since changed, so replay buffers age badly and convergence guarantees from single-agent theory do not carry over.

Environmental randomness and partner change are indistinguishable. An agent has no way to attribute a surprising outcome to the world rather than to its team, so it cannot know whether to keep exploring or to re-learn.

Now the correction, because independent learning is routinely dismissed too quickly.

It is simple. No new algorithm. If you can run PPO, you can run IPPO today.

It scales in the way that matters. Each agent learns over its own action set, so nothing here grows exponentially with the number of agents, which is exactly the wall that centralized joint-action control runs into in the next section.

It is decentralized by construction. There is no training-time infrastructure to dismantle before deployment, and no question about whether the learned policies will run locally. They already do.

Knowledge check

Which statement best describes why non-stationarity is a problem specifically for a learning algorithm?

Select one answer.

  • Independent learning runs a single-agent algorithm per agent, on local observations and local actions, with no explicit model of partners.
  • Partners therefore become part of each agent’s environment, a part that learns, making the perceived transition, observation and reward functions all non-stationary.
  • An agent cannot distinguish environmental randomness from a partner changing policy.
  • Its costs: no use of information about other agents, non-stationarity from concurrent learning, and unattributable surprise.
  • Its strengths: simplicity, scalability in the number of agents, and decentralization by construction.
  • It remains a serious baseline, sometimes competitive with far more sophisticated methods. Treat it as the bar to clear, not the strawman.