ToMATO: Verbalizing the Mental States of Role-Playing LLMs for Benchmarking Theory of Mind
Fuente:
arXiv
Saved in:
| Main Authors: | Shinoda, Kazutoshi, Hojo, Nobukatsu, Nishida, Kyosuke, Mizuno, Saki, Suzuki, Keita, Masumura, Ryo, Sugiyama, Hiroaki, Saito, Kuniko |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Let's Put Ourselves in Sally's Shoes: Shoes-of-Others Prefilling Improves Theory of Mind in Large Language Models
by: Shinoda, Kazutoshi, et al.
Published: (2025)
by: Shinoda, Kazutoshi, et al.
Published: (2025)
Debiasing Reward Models via Causally Motivated Inference-Time Intervention
by: Shinoda, Kazutoshi, et al.
Published: (2026)
by: Shinoda, Kazutoshi, et al.
Published: (2026)
Initialization of Large Language Models via Reparameterization to Mitigate Loss Spikes
by: Nishida, Kosuke, et al.
Published: (2024)
by: Nishida, Kosuke, et al.
Published: (2024)
InstructDoc: A Dataset for Zero-Shot Generalization of Visual Document Understanding with Instructions
by: Tanaka, Ryota, et al.
Published: (2024)
by: Tanaka, Ryota, et al.
Published: (2024)
Wavelet-based Positional Representation for Long Context
by: Oka, Yui, et al.
Published: (2025)
by: Oka, Yui, et al.
Published: (2025)
VDocRAG: Retrieval-Augmented Generation over Visually-Rich Documents
by: Tanaka, Ryota, et al.
Published: (2025)
by: Tanaka, Ryota, et al.
Published: (2025)
Portable Reward Tuning: Towards Reusable Fine-Tuning across Different Pretrained Models
by: Chijiwa, Daiki, et al.
Published: (2025)
by: Chijiwa, Daiki, et al.
Published: (2025)
Can LLMs Detect Their Own Hallucinations?
by: Kadotani, Sora, et al.
Published: (2025)
by: Kadotani, Sora, et al.
Published: (2025)
VoiceGrad: Non-Parallel Any-to-Many Voice Conversion with Annealed Langevin Dynamics
by: Kameoka, Hirokazu, et al.
Published: (2020)
by: Kameoka, Hirokazu, et al.
Published: (2020)
MVAFormer: RGB-based Multi-View Spatio-Temporal Action Recognition with Transformer
by: Yamane, Taiga, et al.
Published: (2025)
by: Yamane, Taiga, et al.
Published: (2025)
MVTrajecter: Multi-View Pedestrian Tracking with Trajectory Motion Cost and Trajectory Appearance Cost
by: Yamane, Taiga, et al.
Published: (2025)
by: Yamane, Taiga, et al.
Published: (2025)
BioVITA: Biological Dataset, Model, and Benchmark for Visual-Textual-Acoustic Alignment
by: Shinoda, Risa, et al.
Published: (2026)
by: Shinoda, Risa, et al.
Published: (2026)
Theory of Mind and Self-Attributions of Mentality are Dissociable in LLMs
by: Kim, Junsol, et al.
Published: (2026)
by: Kim, Junsol, et al.
Published: (2026)
Relative Density Ratio Optimization for Stable and Statistically Consistent Model Alignment
by: Takahashi, Hiroshi, et al.
Published: (2026)
by: Takahashi, Hiroshi, et al.
Published: (2026)
Responses Fall Short of Understanding: Revealing the Gap between Internal Representations and Responses in Visual Document Understanding
by: Kawasaki, Haruka, et al.
Published: (2026)
by: Kawasaki, Haruka, et al.
Published: (2026)
Factor-Conditioned Speaking-Style Captioning
by: Ando, Atsushi, et al.
Published: (2024)
by: Ando, Atsushi, et al.
Published: (2024)
Boundary Potential Method for Describing Electron Teleportation in an Interferometer with a Topological Superconductor
by: Mizuno, Kyosuke, et al.
Published: (2026)
by: Mizuno, Kyosuke, et al.
Published: (2026)
Assessing LLMs in Art Contexts: Critique Generation and Theory of Mind Evaluation
by: Arita, Takaya, et al.
Published: (2025)
by: Arita, Takaya, et al.
Published: (2025)
LLMs and their Limited Theory of Mind: Evaluating Mental State Annotations in Situated Dialogue
by: Kowalyshyn, Katharine, et al.
Published: (2025)
by: Kowalyshyn, Katharine, et al.
Published: (2025)
Reply to the letter
by: Yoshika Saito, et al.
Published: (2026)
by: Yoshika Saito, et al.
Published: (2026)
Innovative Tokyo / Kuniko Fujita, Richard Child Hill
by: Fujita, Kuniko
Published: (2005)
by: Fujita, Kuniko
Published: (2005)
CHARTOM: A Visual Theory-of-Mind Benchmark for LLMs on Misleading Charts
by: Bharti, Shubham, et al.
Published: (2024)
by: Bharti, Shubham, et al.
Published: (2024)
Accurate estimation of measurement position in Brillouin optical correlation-domain reflectometry based on Rayleigh noise spectral analysis
by: Kikuchi, Keita, et al.
Published: (2024)
by: Kikuchi, Keita, et al.
Published: (2024)
User-Specific Dialogue Generation with User Profile-Aware Pre-Training Model and Parameter-Efficient Fine-Tuning
by: Otsuka, Atsushi, et al.
Published: (2024)
by: Otsuka, Atsushi, et al.
Published: (2024)
Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization
by: Kawakami, Wataru, et al.
Published: (2025)
by: Kawakami, Wataru, et al.
Published: (2025)
Rethinking Theory of Mind Benchmarks for LLMs: Towards A User-Centered Perspective
by: Wang, Qiaosi, et al.
Published: (2025)
by: Wang, Qiaosi, et al.
Published: (2025)
On the Role of Hidden States of Modern Hopfield Network in Transformer
by: Masumura, Tsubasa, et al.
Published: (2025)
by: Masumura, Tsubasa, et al.
Published: (2025)
Shifting power, interstate war, and domestic politics
by: Scott Wolford, et al.
Published: (2025)
by: Scott Wolford, et al.
Published: (2025)
Library Aides: Building Character, Advancing Service
by: Muronaga, Karen, et al.
Published: (2008)
by: Muronaga, Karen, et al.
Published: (2008)
Amrita Field Theory (AFT) and Universal Consciousness as Foundational Field: A Comparative Analysis of Layered Information Ontology, Testability, and Scope
by: Nagatsu, Kazutoshi
Published: (2025)
by: Nagatsu, Kazutoshi
Published: (2025)
Theory of Mind Development in Children With Congenital Visual Impairment: Role of Visual Impairment and Verbal Ability
by: Yong Yang, et al.
Published: (2025)
by: Yong Yang, et al.
Published: (2025)
Few-shot Personalization via In-Context Learning for Speech Emotion Recognition based on Speech-Language Model
by: Ihori, Mana, et al.
Published: (2025)
by: Ihori, Mana, et al.
Published: (2025)
MSMVD: Exploiting Multi-scale Image Features via Multi-scale BEV Features for Multi-view Pedestrian Detection
by: Yamane, Taiga, et al.
Published: (2025)
by: Yamane, Taiga, et al.
Published: (2025)
Human Preference-Aligned Concept Customization Benchmark via Decomposed Evaluation
by: Ishikawa, Reina, et al.
Published: (2025)
by: Ishikawa, Reina, et al.
Published: (2025)
Mind the Motions: Benchmarking Theory-of-Mind in Everyday Body Language
by: Lee, Seungbeen, et al.
Published: (2025)
by: Lee, Seungbeen, et al.
Published: (2025)
PetFace: A Large-Scale Dataset and Benchmark for Animal Identification
by: Shinoda, Risa, et al.
Published: (2024)
by: Shinoda, Risa, et al.
Published: (2024)
HalDec-Bench: Benchmarking Hallucination Detector in Image Captioning
by: Saito, Kuniaki, et al.
Published: (2026)
by: Saito, Kuniaki, et al.
Published: (2026)
HalDec-Bench: Benchmarking Hallucination Detector in Image Captioning
by: Saito, Kuniaki, et al.
Published: (2025)
by: Saito, Kuniaki, et al.
Published: (2025)
Benchmarking the Utility of Privacy-Preserving Cox Regression Under Data-Driven Clipping Bounds: A Multi-Dataset Simulation Study
by: Fukuyama, Keita, et al.
Published: (2026)
by: Fukuyama, Keita, et al.
Published: (2026)
GaussianPlant: Structure-aligned Gaussian Splatting for 3D Reconstruction of Plants
by: Yang, Yang, et al.
Published: (2025)
by: Yang, Yang, et al.
Published: (2025)
Similar Items
-
Let's Put Ourselves in Sally's Shoes: Shoes-of-Others Prefilling Improves Theory of Mind in Large Language Models
by: Shinoda, Kazutoshi, et al.
Published: (2025) -
Debiasing Reward Models via Causally Motivated Inference-Time Intervention
by: Shinoda, Kazutoshi, et al.
Published: (2026) -
Initialization of Large Language Models via Reparameterization to Mitigate Loss Spikes
by: Nishida, Kosuke, et al.
Published: (2024) -
InstructDoc: A Dataset for Zero-Shot Generalization of Visual Document Understanding with Instructions
by: Tanaka, Ryota, et al.
Published: (2024) -
Wavelet-based Positional Representation for Long Context
by: Oka, Yui, et al.
Published: (2025)