Guardado en:
Detalles Bibliográficos
Autores principales: Luo, Yifan, Xu, Kangping, Lu, Yanzhen, Yuan, Yang, Yao, Andrew Chi-Chih
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:https://arxiv.org/abs/2603.25187
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911546042482688
author Luo, Yifan
Xu, Kangping
Lu, Yanzhen
Yuan, Yang
Yao, Andrew Chi-Chih
author_facet Luo, Yifan
Xu, Kangping
Lu, Yanzhen
Yuan, Yang
Yao, Andrew Chi-Chih
contents Persona-driven large language models (LLMs) require consistent behavioral tendencies across interactions to simulate human-like personality traits, such as persistence or reliability. However, current LLMs often lack stable internal representations that anchor their responses over extended dialogues. This work explores whether LLMs can maintain "implicit consistency", defined as persistent adherence to an unstated goal in multi-turn interactions. We designed a 20-question-style riddle game paradigm where an LLM is tasked with secretly selecting a target and responding to users' guesses with "yes/no" answers. Through evaluations, we find that LLMs struggle to preserve latent consistency: their implicit "goals" shift across turns unless explicitly provided their selected target in context. These findings highlight critical limitations in the building of persona-driven LLMs and underscore the need for mechanisms that anchor implicit goals over time, which is a key to realistic personality modeling in interactive applications such as dialogue systems.
format Preprint
id arxiv_https___arxiv_org_abs_2603_25187
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Probing the Lack of Stable Internal Beliefs in LLMs
Luo, Yifan
Xu, Kangping
Lu, Yanzhen
Yuan, Yang
Yao, Andrew Chi-Chih
Computation and Language
Artificial Intelligence
Persona-driven large language models (LLMs) require consistent behavioral tendencies across interactions to simulate human-like personality traits, such as persistence or reliability. However, current LLMs often lack stable internal representations that anchor their responses over extended dialogues. This work explores whether LLMs can maintain "implicit consistency", defined as persistent adherence to an unstated goal in multi-turn interactions. We designed a 20-question-style riddle game paradigm where an LLM is tasked with secretly selecting a target and responding to users' guesses with "yes/no" answers. Through evaluations, we find that LLMs struggle to preserve latent consistency: their implicit "goals" shift across turns unless explicitly provided their selected target in context. These findings highlight critical limitations in the building of persona-driven LLMs and underscore the need for mechanisms that anchor implicit goals over time, which is a key to realistic personality modeling in interactive applications such as dialogue systems.
title Probing the Lack of Stable Internal Beliefs in LLMs
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2603.25187