Accumulating Context Changes the Beliefs of Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Geng, Jiayi, Chen, Howard, Liu, Ryan, Ribeiro, Manoel Horta, Willer, Robb, Neubig, Graham, Griffiths, Thomas L.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914135441145856
author Geng, Jiayi
Chen, Howard
Liu, Ryan
Ribeiro, Manoel Horta
Willer, Robb
Neubig, Graham
Griffiths, Thomas L.
author_facet Geng, Jiayi
Chen, Howard
Liu, Ryan
Ribeiro, Manoel Horta
Willer, Robb
Neubig, Graham
Griffiths, Thomas L.
contents Language model (LM) assistants are increasingly used in applications such as brainstorming and research. Improvements in memory and context size have allowed these models to become more autonomous, which has also resulted in more text accumulation in their context windows without explicit user intervention. This comes with a latent risk: the belief profiles of models -- their understanding of the world as manifested in their responses or actions -- may silently change as context accumulates. This can lead to subtly inconsistent user experiences, or shifts in behavior that deviate from the original alignment of the models. In this paper, we explore how accumulating context by engaging in interactions and processing text -- talking and reading -- can change the beliefs of language models, as manifested in their responses and behaviors. Our results reveal that models' belief profiles are highly malleable: GPT-5 exhibits a 54.7% shift in its stated beliefs after 10 rounds of discussion about moral dilemmas and queries about safety, while Grok 4 shows a 27.2% shift on political issues after reading texts from the opposing position. We also examine models' behavioral changes by designing tasks that require tool use, where each tool selection corresponds to an implicit belief. We find that these changes align with stated belief shifts, suggesting that belief shifts will be reflected in actual behavior in agentic systems. Our analysis exposes the hidden risk of belief shift as models undergo extended sessions of talking or reading, rendering their opinions and actions unreliable.
format Preprint
id arxiv_https___arxiv_org_abs_2511_01805
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Accumulating Context Changes the Beliefs of Language Models
Geng, Jiayi
Chen, Howard
Liu, Ryan
Ribeiro, Manoel Horta
Willer, Robb
Neubig, Graham
Griffiths, Thomas L.
Computation and Language
Artificial Intelligence
Language model (LM) assistants are increasingly used in applications such as brainstorming and research. Improvements in memory and context size have allowed these models to become more autonomous, which has also resulted in more text accumulation in their context windows without explicit user intervention. This comes with a latent risk: the belief profiles of models -- their understanding of the world as manifested in their responses or actions -- may silently change as context accumulates. This can lead to subtly inconsistent user experiences, or shifts in behavior that deviate from the original alignment of the models. In this paper, we explore how accumulating context by engaging in interactions and processing text -- talking and reading -- can change the beliefs of language models, as manifested in their responses and behaviors. Our results reveal that models' belief profiles are highly malleable: GPT-5 exhibits a 54.7% shift in its stated beliefs after 10 rounds of discussion about moral dilemmas and queries about safety, while Grok 4 shows a 27.2% shift on political issues after reading texts from the opposing position. We also examine models' behavioral changes by designing tasks that require tool use, where each tool selection corresponds to an implicit belief. We find that these changes align with stated belief shifts, suggesting that belief shifts will be reflected in actual behavior in agentic systems. Our analysis exposes the hidden risk of belief shift as models undergo extended sessions of talking or reading, rendering their opinions and actions unreliable.
title Accumulating Context Changes the Beliefs of Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2511.01805