OSCToM: RL-Guided Adversarial Generation for High-Order Theory of Mind

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Srishty, Sharmin Sultana, Rahman, Kazi Mahathir, Sakkhi, Malaika Parizat, Prianna, Samia Shahid, Sinat, Shaikhul Islam
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914581461336064
author Srishty, Sharmin Sultana
Rahman, Kazi Mahathir
Sakkhi, Malaika Parizat
Prianna, Samia Shahid
Sinat, Shaikhul Islam
author_facet Srishty, Sharmin Sultana
Rahman, Kazi Mahathir
Sakkhi, Malaika Parizat
Prianna, Samia Shahid
Sinat, Shaikhul Islam
contents Large Language Models (LLMs) perform well on many language tasks, but their Theory of Mind (ToM) reasoning is still uneven in complex social settings. Existing benchmarks, including ExploreToM, do not always test the recursive beliefs and information asymmetries that make these settings difficult. This paper presents OSCToM (Observer-Self Conflict Theory of Mind), an approach for modeling nested belief conflicts in LLM-based ToM tasks. The key case is one in which an observer's view of another agent conflicts with the observer's own belief state. Such cases go beyond simple perspective-taking and require recursive, multi-layered reasoning. OSCToM combines reinforcement learning (RL), an extended domain-specific language, and compositional surrogate models to generate observer-self conflicts. In our experiments, OSCToM-8B gives the best overall result among the systems tested. It improves on the reported ExploreToM results on FANToM and remains competitive on Hi-ToM and BigToM. On the information-asymmetric FANToM benchmark, OSCToM reaches 76% accuracy, compared with the 0.2% reported by ExploreToM. The data-synthesis procedure is also 6x more efficient, indicating that targeted training data can help smaller models handle advanced cognitive reasoning. The project code is available at https://github.com/sharminsrishty/osct.
format Preprint
id arxiv_https___arxiv_org_abs_2605_20423
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle OSCToM: RL-Guided Adversarial Generation for High-Order Theory of Mind
Srishty, Sharmin Sultana
Rahman, Kazi Mahathir
Sakkhi, Malaika Parizat
Prianna, Samia Shahid
Sinat, Shaikhul Islam
Artificial Intelligence
I.2.7; I.2.6
Large Language Models (LLMs) perform well on many language tasks, but their Theory of Mind (ToM) reasoning is still uneven in complex social settings. Existing benchmarks, including ExploreToM, do not always test the recursive beliefs and information asymmetries that make these settings difficult. This paper presents OSCToM (Observer-Self Conflict Theory of Mind), an approach for modeling nested belief conflicts in LLM-based ToM tasks. The key case is one in which an observer's view of another agent conflicts with the observer's own belief state. Such cases go beyond simple perspective-taking and require recursive, multi-layered reasoning. OSCToM combines reinforcement learning (RL), an extended domain-specific language, and compositional surrogate models to generate observer-self conflicts. In our experiments, OSCToM-8B gives the best overall result among the systems tested. It improves on the reported ExploreToM results on FANToM and remains competitive on Hi-ToM and BigToM. On the information-asymmetric FANToM benchmark, OSCToM reaches 76% accuracy, compared with the 0.2% reported by ExploreToM. The data-synthesis procedure is also 6x more efficient, indicating that targeted training data can help smaller models handle advanced cognitive reasoning. The project code is available at https://github.com/sharminsrishty/osct.
title OSCToM: RL-Guided Adversarial Generation for High-Order Theory of Mind
topic Artificial Intelligence
I.2.7; I.2.6
url https://arxiv.org/abs/2605.20423