Exploration Through Introspection: A Self-Aware Reward Model

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Petrowski, Michael, Gašić, Milica
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914237743366144
author Petrowski, Michael
Gašić, Milica
author_facet Petrowski, Michael
Gašić, Milica
contents Understanding how artificial agents model internal mental states is central to advancing Theory of Mind in AI. Evidence points to a unified system for self- and other-awareness. We explore this self-awareness by having reinforcement learning agents infer their own internal states in gridworld environments. Specifically, we introduce an introspective exploration component that is inspired by biological pain as a learning signal by utilizing a hidden Markov model to infer "pain-belief" from online observations. This signal is integrated into a subjective reward function to study how self-awareness affects the agent's learning abilities. Further, we use this computational framework to investigate the difference in performance between normal and chronic pain perception models. Results show that introspective agents in general significantly outperform standard baseline agents and can replicate complex human-like behaviors.
format Preprint
id arxiv_https___arxiv_org_abs_2601_03389
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Exploration Through Introspection: A Self-Aware Reward Model
Petrowski, Michael
Gašić, Milica
Artificial Intelligence
Machine Learning
Understanding how artificial agents model internal mental states is central to advancing Theory of Mind in AI. Evidence points to a unified system for self- and other-awareness. We explore this self-awareness by having reinforcement learning agents infer their own internal states in gridworld environments. Specifically, we introduce an introspective exploration component that is inspired by biological pain as a learning signal by utilizing a hidden Markov model to infer "pain-belief" from online observations. This signal is integrated into a subjective reward function to study how self-awareness affects the agent's learning abilities. Further, we use this computational framework to investigate the difference in performance between normal and chronic pain perception models. Results show that introspective agents in general significantly outperform standard baseline agents and can replicate complex human-like behaviors.
title Exploration Through Introspection: A Self-Aware Reward Model
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2601.03389