Skip to content
MARL in Cooperative Environments
Edit this page

3.5Agent Modelling

5 min read

Agent modelling uses observed behaviour to predict how another agent is likely to act.

In this section you will

  • Predict what another agent is likely to do
  • Build an explicit model of a partner’s policy
  • Condition an ego policy on that prediction
  • Reason about what happens when the model is wrong

Suppose agent 1 is paired with an unfamiliar agent, and watches it for a few steps.

StepWhat agent 1 observes the stranger doing
1moves toward the stove
2prepares the cooking station
3stays beside the stove

A reasonable inference:

This agent intends to cook. It probably expects me to collect the ingredients.

And therefore a decision:

I should collect the ingredient.

Nothing exotic happened. Agent 1 used its observations of behaviour to predict behaviour, and acted on the prediction. That is agent modelling, which the textbook defines as constructing models of other agents that make useful predictions about their behaviour.

Formally, agent i\ag maintains an estimate of another agent’s policy.

Figure 1
π^j i\tone{policy}{\hat{\pi}^{\,\ag}_j}
π^j i\hat{\pi}^{\,\ag}_j
agent i’s model OF agent j. The hat marks it as an estimate, and the superscript marks whose estimate it is
A model, held by one agent, of another agent's decision rule.

Two features of that notation carry the content. The hat says this is agent i\ag‘s belief rather than the truth: agent jj‘s actual policy is not available. And the superscript says the model belongs to agent i\ag: two agents modelling the same partner may hold different models of it.

What the model produces is a prediction over the partner’s next action.

Figure 2
Pr⁡(atj∣hti⏟everything agent i has seen so far)\Pr\bigl(\tone{action}{a^{j}_t} \given \ubrace{observe}{h^{\ag}_t}{everything agent i has seen so far}\bigr)
atja^{j}_t
the partner’s next action, which agent i is trying to anticipate
htih^{\ag}_t
agent i’s own history: the only evidence it has
A distribution over what the partner will do, conditioned only on what this agent has personally seen.

Look hard at the right-hand side of that conditioning bar, because it is what keeps this honest. The prediction is conditioned on htih^{\ag}_t, agent i\ag‘s own history. Not on the state, not on the partner’s observations, and certainly not on the partner’s parameters.

The textbook describes exactly this: neural agent models that use an agent’s observation history to predict the actions of other agents.

A prediction is only worth having if it changes a decision. So the model becomes an input to the policy.

Figure 3
ati∼πi(⋅∣oti⏟what I see, π^j i⏟who I think you are)\tone{action}{a^{\ag}_t} \sim \tone{policy}{\pol{\ag}}\bigl(\cdot \given \ubrace{observe}{\obs{\ag}}{what I see},\ \ubrace{policy}{\hat{\pi}^{\,\ag}_j}{who I think you are}\bigr)
oti\obs{\ag}
the observation, as in every previous chapter
π^j i\hat{\pi}^{\,\ag}_j
the partner model, now part of what the policy conditions on
Same local policy as always, with one extra input: a belief about the partner.

In the kitchen this is the difference between two behaviours. Without the model, agent 1 follows a fixed rule and hopes it fits. With it, agent 1 waits two steps, observes a stranger heading for the stove, and chooses to fetch.

Worth stating, because agent modelling is not free.

A model that is correct lets an agent specialise and do better than any partner-agnostic policy could. A model that is confidently wrong makes it do worse: it commits to a complementary role that does not complement, and the failure is decisive rather than hedged.

That is the same asymmetry as the Communication Lab’s cross-play matrix, where mismatched protocols scored 0.00 against a chance baseline of 0.25. Confident misunderstanding beats no information only when the confidence is justified.

Knowledge check

Which of these partner models could actually be used by a deployed agent?

Select one answer.

  • Agent modelling builds models of other agents that make useful predictions about their behaviour.
  • π^j i\hat{\pi}^{\,\ag}_j is agent i\ag‘s estimate of agent jj. Both the hat and the superscript matter.
  • The prediction Pr⁡(atj∣hti)\Pr(a^j_t \given h^\ag_t) is conditioned on the modelling agent’s own history, which is the only evidence it has.
  • A model that takes a partner’s observations, parameters, or the global state is not deployable.
  • The model becomes an input to the policy, so behaviour adapts within an episode without any weights changing.
  • Diversity and modelling are complementary: one produces robustness without identification, the other identifies and specialises.
  • A confidently wrong model is worse than none, because it commits.