3.6Partner Representations
A partner representation is a compact summary of behaviour that helps an agent choose compatible actions without reconstructing another agent’s complete policy.
In this section you will
- Explain why recovering a full policy is usually unnecessary
- Describe a learned behaviour embedding
- Separate behavioural features from identity labels
- Infer a representation online as evidence arrives
Limits of Full Policy Recovery
Section titled “Limits of Full Policy Recovery”Three obstacles, and the textbook names all of them as motivation for a compact representation instead. Other agents’ policies may be
- unknown at execution, since nothing gives you access to them,
- too complex to represent directly, being large neural networks,
- changing over time, so any reconstruction goes stale.
The third is the sharpest. Even a perfect reconstruction of a partner’s policy describes the partner as it was, and a partner that is still learning has moved on. This is the non-stationarity of the Coordinate chapter, arriving here as an obsolescence problem.
Behaviour Embeddings
Section titled “Behaviour Embeddings”So do not reconstruct the policy. Learn a summary that is useful for interacting with it.
- the modelling agent’s own observation history
- a learned vector standing for agent j’s behaviour
Then condition on that instead.
- agent i’s selected action
- agent i’s local observation
- the inferred representation of partner j
The gain is that can be much smaller than a policy and still carry what matters. Deciding whether to fetch or cook needs to know roughly what kind of partner this is, not how it would behave in every situation it will never encounter.
Behaviour Rather Than Identity
Section titled “Behaviour Rather Than Identity”Here is the distinction that decides whether a representation is any use.
| What is learned | What it supports | |
|---|---|---|
| Identity | partner = #7 do X | recognising partners seen in training |
| Behaviour | prefers cooking, commits early, responds to requests, takes left-side tasks | acting sensibly with a partner never seen before |
An identity representation is a lookup table over training partners. It can be learned, it can score well on any evaluation that reuses those partners, and it fails completely on a stranger, because a stranger has no entry.
A behaviour representation places a partner in a space of behaviours. A new partner that happens to behave like a known one lands nearby and inherits a sensible response. This is the same argument as the difference between memorising and generalising in supervised learning, applied to partners.
Online Representation Inference
Section titled “Online Representation Inference”The representation is not fixed at the start of an episode. It is revised as evidence arrives.
- the interaction history available at step t
- the partner representation inferred from that history
Early in an episode the agent has seen almost nothing, so should be close to uninformative and the agent’s behaviour appropriately non-committal. After several steps of watching a partner head for the stove every time, is much more specific and the agent can commit.
It also gives a shape worth expecting in results: performance that starts near partner-agnostic and improves over the course of an episode. An adaptive agent that shows no such curve is probably not using its model.
Knowledge check
Correct.
Not quite.
The encoder learned partner identity rather than partner behaviour, so a stranger falls outside the five categories it can represent.
Exactly. Five vectors for five training partners is a lookup table. It scores well on the partners it was built from and has nothing to say about a sixth, which is why the advantage vanishes on held-out partners.
The encoder needs more training on the five partners.
More training on the same five would sharpen the lookup table, making the identity shortcut stronger rather than weaker. The problem is what it is representing, not how well.
Five partners is too few to learn from.
A contributing factor and not the mechanism. Even with five genuinely different partners, an encoder that maps to five discrete identities cannot place a sixth. The fix is a representation of behaviour, not just more of it.
Held-out partners are always harder, so this is expected.
Held-out partners are harder, and the specific finding here is that the model provided no benefit at all over ignoring the partner. That is the signature of a representation that does not transfer.
Explanation
“Would this representation say anything about a partner it has never seen?” is the question that separates the two cases.
Partner Representations Summary
Section titled “Partner Representations Summary”- Reconstructing a partner’s policy is hard: policies are unknown at execution, too complex to represent directly, and changing over time.
- A behaviour embedding summarises how a partner behaves, learned from the modelling agent’s own history, and only as detailed as acting well requires.
- Representing behaviour generalises to strangers. Representing identity is a lookup table over training partners and does not.
- Identity shortcuts are invisible to any evaluation that reuses training partners.
- is revised online as evidence arrives, which is adaptation without any weight update. Expect within-episode improvement as the signature.