Indications of Belief-Guided Agency and Meta-Cognitive Monitoring in Large Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Yalon, Noam Steinmetz, Goldstein, Ariel, Mudrik, Liad, Geva, Mor
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917242499760128
author Yalon, Noam Steinmetz
Goldstein, Ariel
Mudrik, Liad
Geva, Mor
author_facet Yalon, Noam Steinmetz
Goldstein, Ariel
Mudrik, Liad
Geva, Mor
contents Rapid advancements in large language models (LLMs) have sparked the question whether these models possess some form of consciousness. To tackle this challenge, Butlin et al. (2023) introduced a list of indicators for consciousness in artificial systems based on neuroscientific theories. In this work, we evaluate a key indicator from this list, called HOT-3, which tests for agency guided by a general belief-formation and action selection system that updates beliefs based on meta-cognitive monitoring. We view beliefs as representations in the model's latent space that emerge in response to a given input, and introduce a metric to quantify their dominance during generation. Analyzing the dynamics between competing beliefs across models and tasks reveals three key findings: (1) external manipulations systematically modulate internal belief formation, (2) belief formation causally drives the model's action selection, and (3) models can monitor and report their own belief states. Together, these results provide empirical support for the existence of belief-guided agency and meta-cognitive monitoring in LLMs. More broadly, our work lays methodological groundwork for investigating the emergence of agency, beliefs, and meta-cognition in LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2602_02467
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Indications of Belief-Guided Agency and Meta-Cognitive Monitoring in Large Language Models
Yalon, Noam Steinmetz
Goldstein, Ariel
Mudrik, Liad
Geva, Mor
Computation and Language
Rapid advancements in large language models (LLMs) have sparked the question whether these models possess some form of consciousness. To tackle this challenge, Butlin et al. (2023) introduced a list of indicators for consciousness in artificial systems based on neuroscientific theories. In this work, we evaluate a key indicator from this list, called HOT-3, which tests for agency guided by a general belief-formation and action selection system that updates beliefs based on meta-cognitive monitoring. We view beliefs as representations in the model's latent space that emerge in response to a given input, and introduce a metric to quantify their dominance during generation. Analyzing the dynamics between competing beliefs across models and tasks reveals three key findings: (1) external manipulations systematically modulate internal belief formation, (2) belief formation causally drives the model's action selection, and (3) models can monitor and report their own belief states. Together, these results provide empirical support for the existence of belief-guided agency and meta-cognitive monitoring in LLMs. More broadly, our work lays methodological groundwork for investigating the emergence of agency, beliefs, and meta-cognition in LLMs.
title Indications of Belief-Guided Agency and Meta-Cognitive Monitoring in Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2602.02467