Reasoning Promotes Robustness in Theory of Mind Tasks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: de Haan, Ian B., van der Putten, Peter, van Duijn, Max
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911394047197184
author de Haan, Ian B.
van der Putten, Peter
van Duijn, Max
author_facet de Haan, Ian B.
van der Putten, Peter
van Duijn, Max
contents Large language models (LLMs) have recently shown strong performance on Theory of Mind (ToM) tests, prompting debate about the nature and true performance of the underlying capabilities. At the same time, reasoning-oriented LLMs trained via reinforcement learning with verifiable rewards (RLVR) have achieved notable improvements across a range of benchmarks. This paper examines the behavior of such reasoning models in ToM tasks, using novel adaptations of machine psychological experiments and results from established benchmarks. We observe that reasoning models consistently exhibit increased robustness to prompt variations and task perturbations. Our analysis indicates that the observed gains are more plausibly attributed to increased robustness in finding the correct solution, rather than to fundamentally new forms of ToM reasoning. We discuss the implications of this interpretation for evaluating social-cognitive behavior in LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2601_16853
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Reasoning Promotes Robustness in Theory of Mind Tasks
de Haan, Ian B.
van der Putten, Peter
van Duijn, Max
Artificial Intelligence
Computation and Language
68T50 (Primary), 68T01 (Secondary)
I.2.7; I.2.0
Large language models (LLMs) have recently shown strong performance on Theory of Mind (ToM) tests, prompting debate about the nature and true performance of the underlying capabilities. At the same time, reasoning-oriented LLMs trained via reinforcement learning with verifiable rewards (RLVR) have achieved notable improvements across a range of benchmarks. This paper examines the behavior of such reasoning models in ToM tasks, using novel adaptations of machine psychological experiments and results from established benchmarks. We observe that reasoning models consistently exhibit increased robustness to prompt variations and task perturbations. Our analysis indicates that the observed gains are more plausibly attributed to increased robustness in finding the correct solution, rather than to fundamentally new forms of ToM reasoning. We discuss the implications of this interpretation for evaluating social-cognitive behavior in LLMs.
title Reasoning Promotes Robustness in Theory of Mind Tasks
topic Artificial Intelligence
Computation and Language
68T50 (Primary), 68T01 (Secondary)
I.2.7; I.2.0
url https://arxiv.org/abs/2601.16853