LLMs and their Limited Theory of Mind: Evaluating Mental State Annotations in Situated Dialogue
Fuente:
arXiv
Saved in:
| Main Authors: | Kowalyshyn, Katharine, Scheutz, Matthias |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Are you with me? A Framework for Detecting Mental Model Discrepancies in Task-Based Team Dialogues
by: Kowalyshyn, Katharine, et al.
Published: (2026)
by: Kowalyshyn, Katharine, et al.
Published: (2026)
IntelliProof: An Argumentation Network-based Conversational Helper for Organized Reflection
by: Miandoab, Kaveh Eskandari, et al.
Published: (2025)
by: Miandoab, Kaveh Eskandari, et al.
Published: (2025)
MindDial: Belief Dynamics Tracking with Theory-of-Mind Modeling for Situated Neural Dialogue Generation
by: Qiu, Shuwen, et al.
Published: (2023)
by: Qiu, Shuwen, et al.
Published: (2023)
Theory of Mind and Self-Attributions of Mentality are Dissociable in LLMs
by: Kim, Junsol, et al.
Published: (2026)
by: Kim, Junsol, et al.
Published: (2026)
ToMATO: Verbalizing the Mental States of Role-Playing LLMs for Benchmarking Theory of Mind
by: Shinoda, Kazutoshi, et al.
Published: (2025)
by: Shinoda, Kazutoshi, et al.
Published: (2025)
ToM-SSI: Evaluating Theory of Mind in Situated Social Interactions
by: Bortoletto, Matteo, et al.
Published: (2025)
by: Bortoletto, Matteo, et al.
Published: (2025)
Where Norms and References Collide: Evaluating LLMs on Normative Reasoning
by: Abrams, Mitchell, et al.
Published: (2026)
by: Abrams, Mitchell, et al.
Published: (2026)
PersuasiveToM: A Benchmark for Evaluating Machine Theory of Mind in Persuasive Dialogues
by: Yu, Fangxu, et al.
Published: (2025)
by: Yu, Fangxu, et al.
Published: (2025)
Using Machine Mental Imagery for Representing Common Ground in Situated Dialogue
by: Mohapatra, Biswesh, et al.
Published: (2026)
by: Mohapatra, Biswesh, et al.
Published: (2026)
Exploring Safety Alignment Evaluation of LLMs in Chinese Mental Health Dialogues via LLM-as-Judge
by: Cai, Yunna, et al.
Published: (2025)
by: Cai, Yunna, et al.
Published: (2025)
MIST: Towards Multi-dimensional Implicit BiaS Evaluation of LLMs for Theory of Mind
by: Li, Yanlin, et al.
Published: (2025)
by: Li, Yanlin, et al.
Published: (2025)
Assessing LLMs in Art Contexts: Critique Generation and Theory of Mind Evaluation
by: Arita, Takaya, et al.
Published: (2025)
by: Arita, Takaya, et al.
Published: (2025)
RecMind: Japanese Movie Recommendation Dialogue with Seeker's Internal State
by: Kodama, Takashi, et al.
Published: (2024)
by: Kodama, Takashi, et al.
Published: (2024)
Towards A Holistic Landscape of Situated Theory of Mind in Large Language Models
by: Ma, Ziqiao, et al.
Published: (2023)
by: Ma, Ziqiao, et al.
Published: (2023)
On the Benchmarking of LLMs for Open-Domain Dialogue Evaluation
by: Mendonça, John, et al.
Published: (2024)
by: Mendonça, John, et al.
Published: (2024)
MindCube: Spatial Mental Modeling from Limited Views
by: Wang, Qineng, et al.
Published: (2025)
by: Wang, Qineng, et al.
Published: (2025)
EnigmaToM: Improve LLMs' Theory-of-Mind Reasoning Capabilities with Neural Knowledge Base of Entity States
by: Xu, Hainiu, et al.
Published: (2025)
by: Xu, Hainiu, et al.
Published: (2025)
Do LLMs Exhibit Human-Like Reasoning? Evaluating Theory of Mind in LLMs for Open-Ended Responses
by: Amirizaniani, Maryam, et al.
Published: (2024)
by: Amirizaniani, Maryam, et al.
Published: (2024)
Towards Dynamic Theory of Mind: Evaluating LLM Adaptation to Temporal Evolution of Human States
by: Xiao, Yang, et al.
Published: (2025)
by: Xiao, Yang, et al.
Published: (2025)
Noise Injection Systemically Degrades Large Language Model Safety Guardrails
by: Shahani, Prithviraj Singh, et al.
Published: (2025)
by: Shahani, Prithviraj Singh, et al.
Published: (2025)
Mind the Quote: Enabling Quotation-Aware Dialogue in LLMs via Plug-and-Play Modules
by: Zhang, Yueqi, et al.
Published: (2025)
by: Zhang, Yueqi, et al.
Published: (2025)
Large Language Model based Situational Dialogues for Second Language Learning
by: Xu, Shuyao, et al.
Published: (2024)
by: Xu, Shuyao, et al.
Published: (2024)
OMIND: Framework for Knowledge Grounded Finetuning and Multi-Turn Dialogue Benchmark for Mental Health LLMs
by: Racha, Suraj, et al.
Published: (2026)
by: Racha, Suraj, et al.
Published: (2026)
An Annotation Scheme and Classifier for Personal Facts in Dialogue
by: Zaitsev, Konstantin
Published: (2026)
by: Zaitsev, Konstantin
Published: (2026)
Designing and Evaluating Dialogue LLMs for Co-Creative Improvised Theatre
by: Branch, Boyd, et al.
Published: (2024)
by: Branch, Boyd, et al.
Published: (2024)
Soda-Eval: Open-Domain Dialogue Evaluation in the age of LLMs
by: Mendonça, John, et al.
Published: (2024)
by: Mendonça, John, et al.
Published: (2024)
Persuasion Should be Double-Blind: A Multi-Domain Dialogue Dataset With Faithfulness Based on Causal Theory of Mind
by: Zhang, Dingyi, et al.
Published: (2025)
by: Zhang, Dingyi, et al.
Published: (2025)
TRACE: Real-Time Multimodal Common Ground Tracking in Situated Collaborative Dialogues
by: VanderHoeven, Hannah, et al.
Published: (2025)
by: VanderHoeven, Hannah, et al.
Published: (2025)
Are AI Machines Making Humans Obsolete?
by: Scheutz, Matthias
Published: (2025)
by: Scheutz, Matthias
Published: (2025)
CoMMET: To What Extent Can LLMs Perform Theory of Mind Tasks?
by: Chen, Ruirui, et al.
Published: (2026)
by: Chen, Ruirui, et al.
Published: (2026)
Multi-Turn Puzzles: Evaluating Interactive Reasoning and Strategic Dialogue in LLMs
by: Badola, Kartikeya, et al.
Published: (2025)
by: Badola, Kartikeya, et al.
Published: (2025)
Are they lovers or friends? Evaluating LLMs' Social Reasoning in English and Korean Dialogues
by: Kim, Eunsu, et al.
Published: (2025)
by: Kim, Eunsu, et al.
Published: (2025)
Are LLMs Robust for Spoken Dialogues?
by: Mousavi, Seyed Mahed, et al.
Published: (2024)
by: Mousavi, Seyed Mahed, et al.
Published: (2024)
Interpretable and Robust Dialogue State Tracking via Natural Language Summarization with LLMs
by: Carranza, Rafael, et al.
Published: (2025)
by: Carranza, Rafael, et al.
Published: (2025)
Evaluating Large Language Models in Theory of Mind Tasks
by: Kosinski, Michal
Published: (2023)
by: Kosinski, Michal
Published: (2023)
Are LLMs Effective Negotiators? Systematic Evaluation of the Multifaceted Capabilities of LLMs in Negotiation Dialogues
by: Kwon, Deuksin, et al.
Published: (2024)
by: Kwon, Deuksin, et al.
Published: (2024)
MentalQA: An Annotated Arabic Corpus for Questions and Answers of Mental Healthcare
by: Alhuzali, Hassan, et al.
Published: (2024)
by: Alhuzali, Hassan, et al.
Published: (2024)
A Structured Framework for Evaluating and Enhancing Interpretive Capabilities of Multimodal LLMs in Culturally Situated Tasks
by: Yu, Haorui, et al.
Published: (2025)
by: Yu, Haorui, et al.
Published: (2025)
MEDAL: A Framework for Benchmarking LLMs as Multilingual Open-Domain Dialogue Evaluators
by: Mendonça, John, et al.
Published: (2025)
by: Mendonça, John, et al.
Published: (2025)
UniToMBench: Integrating Perspective-Taking to Improve Theory of Mind in LLMs
by: Thiyagarajan, Prameshwar, et al.
Published: (2025)
by: Thiyagarajan, Prameshwar, et al.
Published: (2025)
Similar Items
-
Are you with me? A Framework for Detecting Mental Model Discrepancies in Task-Based Team Dialogues
by: Kowalyshyn, Katharine, et al.
Published: (2026) -
IntelliProof: An Argumentation Network-based Conversational Helper for Organized Reflection
by: Miandoab, Kaveh Eskandari, et al.
Published: (2025) -
MindDial: Belief Dynamics Tracking with Theory-of-Mind Modeling for Situated Neural Dialogue Generation
by: Qiu, Shuwen, et al.
Published: (2023) -
Theory of Mind and Self-Attributions of Mentality are Dissociable in LLMs
by: Kim, Junsol, et al.
Published: (2026) -
ToMATO: Verbalizing the Mental States of Role-Playing LLMs for Benchmarking Theory of Mind
by: Shinoda, Kazutoshi, et al.
Published: (2025)