When Should Models Change Their Minds? Contextual Belief Management in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Haoming, Xu, Weihong, Li, Zongrui, Wang, Mengru, Yao, Yunzhi, Wu, Chiyu, Shang, Jin, Gong, Yu, Deng, Shumin
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911728612147200
author Xu, Haoming
Xu, Weihong
Li, Zongrui
Wang, Mengru
Yao, Yunzhi
Wu, Chiyu
Shang, Jin
Gong, Yu
Deng, Shumin
author_facet Xu, Haoming
Xu, Weihong
Li, Zongrui
Wang, Mengru
Yao, Yunzhi
Wu, Chiyu
Shang, Jin
Gong, Yu
Deng, Shumin
contents Long-horizon interactions require language models to manage accumulating information: when to update their state, when to preserve their state, and what to ignore. We study this challenge as \textbf{Contextual Belief Management (CBM)}: maintaining a predicted belief state aligned with formal evidence while isolating task-irrelevant noise. To make CBM measurable, we introduce BeliefTrack, a closed-world benchmark spanning Rule Discovery and Circuit Diagnosis, where a finite belief space and symbolic verifiers enable exact turn-level evaluation. BeliefTrack diagnoses three failures: Failed Stay, Failed Update, and Failed Isolation. Across multiple LLMs, vanilla models exhibit severe CBM failures, while explicit belief-tracking prompts provide limited gains. In contrast, reinforcement learning with belief-state rewards reduces failure rates by 70.9\% on average. Further probing reveals latent belief-state dynamics behind these failures, and representation-level steering reduces failure rates by 46.1\% across two tasks\footnote{Code is coming soon at https://github.com/zjunlp/CBM.
format Preprint
id arxiv_https___arxiv_org_abs_2605_30219
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle When Should Models Change Their Minds? Contextual Belief Management in Large Language Models
Xu, Haoming
Xu, Weihong
Li, Zongrui
Wang, Mengru
Yao, Yunzhi
Wu, Chiyu
Shang, Jin
Gong, Yu
Deng, Shumin
Artificial Intelligence
Computation and Language
Machine Learning
Long-horizon interactions require language models to manage accumulating information: when to update their state, when to preserve their state, and what to ignore. We study this challenge as \textbf{Contextual Belief Management (CBM)}: maintaining a predicted belief state aligned with formal evidence while isolating task-irrelevant noise. To make CBM measurable, we introduce BeliefTrack, a closed-world benchmark spanning Rule Discovery and Circuit Diagnosis, where a finite belief space and symbolic verifiers enable exact turn-level evaluation. BeliefTrack diagnoses three failures: Failed Stay, Failed Update, and Failed Isolation. Across multiple LLMs, vanilla models exhibit severe CBM failures, while explicit belief-tracking prompts provide limited gains. In contrast, reinforcement learning with belief-state rewards reduces failure rates by 70.9\% on average. Further probing reveals latent belief-state dynamics behind these failures, and representation-level steering reduces failure rates by 46.1\% across two tasks\footnote{Code is coming soon at https://github.com/zjunlp/CBM.
title When Should Models Change Their Minds? Contextual Belief Management in Large Language Models
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2605.30219