Language Models use Lookbacks to Track Beliefs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Prakash, Nikhil, Shapira, Natalie, Sharma, Arnab Sen, Riedl, Christoph, Belinkov, Yonatan, Shaham, Tamar Rott, Bau, David, Geiger, Atticus
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915814351831040
author Prakash, Nikhil
Shapira, Natalie
Sharma, Arnab Sen
Riedl, Christoph
Belinkov, Yonatan
Shaham, Tamar Rott
Bau, David
Geiger, Atticus
author_facet Prakash, Nikhil
Shapira, Natalie
Sharma, Arnab Sen
Riedl, Christoph
Belinkov, Yonatan
Shaham, Tamar Rott
Bau, David
Geiger, Atticus
contents How do language models (LMs) represent characters' beliefs, especially when those beliefs may differ from reality? This question lies at the heart of understanding the Theory of Mind (ToM) capabilities of LMs. We analyze LMs' ability to reason about characters' beliefs using causal mediation and abstraction. We construct a dataset, CausalToM, consisting of simple stories where two characters independently change the state of two objects, potentially unaware of each other's actions. Our investigation uncovers a pervasive algorithmic pattern that we call a lookback mechanism, which enables the LM to recall important information when it becomes necessary. The LM binds each character-object-state triple together by co-locating their reference information, represented as Ordering IDs (OIs), in low-rank subspaces of the state token's residual stream. When asked about a character's beliefs regarding the state of an object, the binding lookback retrieves the correct state OI and then the answer lookback retrieves the corresponding state token. When we introduce text specifying that one character is (not) visible to the other, we find that the LM first generates a visibility ID encoding the relation between the observing and the observed character OIs. In a visibility lookback, this ID is used to retrieve information about the observed character and update the observing character's beliefs. Our work provides insights into belief tracking mechanisms, taking a step toward reverse-engineering ToM reasoning in LMs.
format Preprint
id arxiv_https___arxiv_org_abs_2505_14685
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Language Models use Lookbacks to Track Beliefs
Prakash, Nikhil
Shapira, Natalie
Sharma, Arnab Sen
Riedl, Christoph
Belinkov, Yonatan
Shaham, Tamar Rott
Bau, David
Geiger, Atticus
Computation and Language
How do language models (LMs) represent characters' beliefs, especially when those beliefs may differ from reality? This question lies at the heart of understanding the Theory of Mind (ToM) capabilities of LMs. We analyze LMs' ability to reason about characters' beliefs using causal mediation and abstraction. We construct a dataset, CausalToM, consisting of simple stories where two characters independently change the state of two objects, potentially unaware of each other's actions. Our investigation uncovers a pervasive algorithmic pattern that we call a lookback mechanism, which enables the LM to recall important information when it becomes necessary. The LM binds each character-object-state triple together by co-locating their reference information, represented as Ordering IDs (OIs), in low-rank subspaces of the state token's residual stream. When asked about a character's beliefs regarding the state of an object, the binding lookback retrieves the correct state OI and then the answer lookback retrieves the corresponding state token. When we introduce text specifying that one character is (not) visible to the other, we find that the LM first generates a visibility ID encoding the relation between the observing and the observed character OIs. In a visibility lookback, this ID is used to retrieve information about the observed character and update the observing character's beliefs. Our work provides insights into belief tracking mechanisms, taking a step toward reverse-engineering ToM reasoning in LMs.
title Language Models use Lookbacks to Track Beliefs
topic Computation and Language
url https://arxiv.org/abs/2505.14685