3.5Agent Modelling
Agent modelling uses observed behaviour to predict how another agent is likely to act.
In this section you will
- Predict what another agent is likely to do
- Build an explicit model of a partner’s policy
- Condition an ego policy on that prediction
- Reason about what happens when the model is wrong
Predicting Other Agents
Section titled “Predicting Other Agents”Suppose agent 1 is paired with an unfamiliar agent, and watches it for a few steps.
| Step | What agent 1 observes the stranger doing |
|---|---|
| 1 | moves toward the stove |
| 2 | prepares the cooking station |
| 3 | stays beside the stove |
A reasonable inference:
This agent intends to cook. It probably expects me to collect the ingredients.
And therefore a decision:
I should collect the ingredient.
Nothing exotic happened. Agent 1 used its observations of behaviour to predict behaviour, and acted on the prediction. That is agent modelling, which the textbook defines as constructing models of other agents that make useful predictions about their behaviour.
Partner Policy Models
Section titled “Partner Policy Models”Formally, agent maintains an estimate of another agent’s policy.
- agent i’s model OF agent j. The hat marks it as an estimate, and the superscript marks whose estimate it is
Two features of that notation carry the content. The hat says this is agent ‘s belief rather than the truth: agent ‘s actual policy is not available. And the superscript says the model belongs to agent : two agents modelling the same partner may hold different models of it.
What the model produces is a prediction over the partner’s next action.
- the partner’s next action, which agent i is trying to anticipate
- agent i’s own history: the only evidence it has
Look hard at the right-hand side of that conditioning bar, because it is what keeps this honest. The prediction is conditioned on , agent ‘s own history. Not on the state, not on the partner’s observations, and certainly not on the partner’s parameters.
The textbook describes exactly this: neural agent models that use an agent’s observation history to predict the actions of other agents.
Partner-Conditioned Policies
Section titled “Partner-Conditioned Policies”A prediction is only worth having if it changes a decision. So the model becomes an input to the policy.
- the observation, as in every previous chapter
- the partner model, now part of what the policy conditions on
In the kitchen this is the difference between two behaviours. Without the model, agent 1 follows a fixed rule and hopes it fits. With it, agent 1 waits two steps, observes a stranger heading for the stove, and chooses to fetch.
Consequences of Model Error
Section titled “Consequences of Model Error”Worth stating, because agent modelling is not free.
A model that is correct lets an agent specialise and do better than any partner-agnostic policy could. A model that is confidently wrong makes it do worse: it commits to a complementary role that does not complement, and the failure is decisive rather than hedged.
That is the same asymmetry as the Communication Lab’s cross-play matrix, where mismatched protocols scored 0.00 against a chance baseline of 0.25. Confident misunderstanding beats no information only when the confidence is justified.
Knowledge check
Correct.
Not quite.
A model predicting the partner’s next action from the modelling agent’s own observation history.
Correct. Its only input is h_i, which the agent has by definition. This is the form the textbook describes, and it is the only one on this list that survives deployment.
A model that takes the partner’s current observation and predicts its action.
The partner’s observation is not available: that is the premise of partial observability. This model could be trained under CTDE and would need distilling into something local before deployment.
A model that reads the partner’s policy parameters directly.
This assumes access to another agent’s internals, which ad hoc teamwork specifically rules out. If you had the parameters you would not need a model.
A model conditioned on the global state, since the state determines everything.
The state does determine a great deal and no agent receives it. This is the same mistake as writing a policy on the state instead of the observation, one level up.
Explanation
“Is every input in ?” is the check that keeps a partner model deployable.
Agent Modelling Summary
Section titled “Agent Modelling Summary”- Agent modelling builds models of other agents that make useful predictions about their behaviour.
- is agent ‘s estimate of agent . Both the hat and the superscript matter.
- The prediction is conditioned on the modelling agent’s own history, which is the only evidence it has.
- A model that takes a partner’s observations, parameters, or the global state is not deployable.
- The model becomes an input to the policy, so behaviour adapts within an episode without any weights changing.
- Diversity and modelling are complementary: one produces robustness without identification, the other identifies and specialises.
- A confidently wrong model is worse than none, because it commits.