Skip to content
MARL in Cooperative Environments
Edit this page

3.6Partner Representations

5 min read

A partner representation is a compact summary of behaviour that helps an agent choose compatible actions without reconstructing another agent’s complete policy.

In this section you will

  • Explain why recovering a full policy is usually unnecessary
  • Describe a learned behaviour embedding
  • Separate behavioural features from identity labels
  • Infer a representation online as evidence arrives

Three obstacles, and the textbook names all of them as motivation for a compact representation instead. Other agents’ policies may be

  • unknown at execution, since nothing gives you access to them,
  • too complex to represent directly, being large neural networks,
  • changing over time, so any reconstruction goes stale.

The third is the sharpest. Even a perfect reconstruction of a partner’s policy describes the partner as it was, and a partner that is still learning has moved on. This is the non-stationarity of the Coordinate chapter, arriving here as an obsolescence problem.

So do not reconstruct the policy. Learn a summary that is useful for interacting with it.

Figure 1
hti⏟what I have seen of you  →  encoder    zj⏟a compact summary of how you behave\ubrace{observe}{h^{\ag}_t}{what I have seen of you} \;\xrightarrow{\;\text{encoder}\;}\; \ubrace{policy}{z_j}{a compact summary of how you behave}
htih^{\ag}_t
the modelling agent’s own observation history
zjz_j
a learned vector standing for agent j’s behaviour
Not a copy of the partner's policy. A description of it, only as detailed as acting well requires.

Then condition on that instead.

Figure 2
ati∼πi(⋅∣oti, zj)\tone{action}{a^{\ag}_t} \sim \tone{policy}{\pol{\ag}}\bigl(\cdot \given \tone{observe}{\obs{\ag}},\ \tone{policy}{z_j}\bigr)
atia^i_t
agent i’s selected action
otio^i_t
agent i’s local observation
zjz_j
the inferred representation of partner j
The local policy conditions its action on both observation and inferred partner behaviour.

The gain is that zjz_j can be much smaller than a policy and still carry what matters. Deciding whether to fetch or cook needs to know roughly what kind of partner this is, not how it would behave in every situation it will never encounter.

Here is the distinction that decides whether a representation is any use.

What is learnedWhat it supports
Identitypartner = #7 →\rightarrow do Xrecognising partners seen in training
Behaviourprefers cooking, commits early, responds to requests, takes left-side tasksacting sensibly with a partner never seen before

An identity representation is a lookup table over training partners. It can be learned, it can score well on any evaluation that reuses those partners, and it fails completely on a stranger, because a stranger has no entry.

A behaviour representation places a partner in a space of behaviours. A new partner that happens to behave like a known one lands nearby and inherits a sensible response. This is the same argument as the difference between memorising and generalising in supervised learning, applied to partners.

The representation is not fixed at the start of an episode. It is revised as evidence arrives.

Figure 3
h1→z1h2→z2h3→z3⋯\tone{observe}{h_1} \rightarrow \tone{policy}{z_1} \qquad \tone{observe}{h_2} \rightarrow \tone{policy}{z_2} \qquad \tone{observe}{h_3} \rightarrow \tone{policy}{z_3} \qquad \cdots
hth_t
the interaction history available at step t
ztz_t
the partner representation inferred from that history
The partner representation is updated as interaction history accumulates.

Early in an episode the agent has seen almost nothing, so zz should be close to uninformative and the agent’s behaviour appropriately non-committal. After several steps of watching a partner head for the stove every time, zz is much more specific and the agent can commit.

It also gives a shape worth expecting in results: performance that starts near partner-agnostic and improves over the course of an episode. An adaptive agent that shows no such curve is probably not using its model.

Knowledge check

An agent's partner encoder maps observation histories to one of five learned vectors, each corresponding to one training partner. On held-out partners it performs no better than a partner-agnostic policy. What is the most likely explanation?

Select one answer.

  • Reconstructing a partner’s policy is hard: policies are unknown at execution, too complex to represent directly, and changing over time.
  • A behaviour embedding zjz_j summarises how a partner behaves, learned from the modelling agent’s own history, and only as detailed as acting well requires.
  • Representing behaviour generalises to strangers. Representing identity is a lookup table over training partners and does not.
  • Identity shortcuts are invisible to any evaluation that reuses training partners.
  • zz is revised online as evidence arrives, which is adaptation without any weight update. Expect within-episode improvement as the signature.