Rethinking Theory of Mind Benchmarks for LLMs: Towards A User-Centered Perspective
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Qiaosi, Zhou, Xuhui, Sap, Maarten, Forlizzi, Jodi, Shen, Hong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Situated, Dynamic, and Subjective: Envisioning the Design of Theory-of-Mind-Enabled Everyday AI with Industry Practitioners
by: Wang, Qiaosi, et al.
Published: (2026)
by: Wang, Qiaosi, et al.
Published: (2026)
Exploring the Innovation Opportunities for Pre-trained Models
by: Park, Minjung, et al.
Published: (2025)
by: Park, Minjung, et al.
Published: (2025)
Minion: A Technology Probe to Explore How Users Negotiate Harmful Value Conflicts with AI Companions
by: Fan, Xianzhe, et al.
Published: (2024)
by: Fan, Xianzhe, et al.
Published: (2024)
Mutual Theory of Mind for Human-AI Communication
by: Wang, Qiaosi, et al.
Published: (2022)
by: Wang, Qiaosi, et al.
Published: (2022)
Promoting Critical Thinking With Domain-Specific Generative AI Provocations
by: von Davier, Thomas Şerban, et al.
Published: (2026)
by: von Davier, Thomas Şerban, et al.
Published: (2026)
AI Mismatches: Identifying Potential Algorithmic Harms Before AI Development
by: Saxena, Devansh, et al.
Published: (2025)
by: Saxena, Devansh, et al.
Published: (2025)
Rel-A.I.: An Interaction-Centered Approach To Measuring Human-LM Reliance
by: Zhou, Kaitlyn, et al.
Published: (2024)
by: Zhou, Kaitlyn, et al.
Published: (2024)
User-Driven Value Alignment: Understanding Users' Perceptions and Strategies for Addressing Biased and Discriminatory Statements in AI Companions
by: Fan, Xianzhe, et al.
Published: (2024)
by: Fan, Xianzhe, et al.
Published: (2024)
Imperfectly Cooperative Human-AI Interactions: Comparing the Impacts of Human and AI Attributes in Simulated and User Studies
by: Cohen, Myke C., et al.
Published: (2026)
by: Cohen, Myke C., et al.
Published: (2026)
Can You Keep a Secret? Exploring AI for Care Coordination in Cognitive Decline
by: Alicia, et al.
Published: (2025)
by: Alicia, et al.
Published: (2025)
Relying on the Unreliable: The Impact of Language Models' Reluctance to Express Uncertainty
by: Zhou, Kaitlyn, et al.
Published: (2024)
by: Zhou, Kaitlyn, et al.
Published: (2024)
Let Them Down Easy! Contextual Effects of LLM Guardrails on User Perceptions and Preferences
by: Zheng, Mingqian, et al.
Published: (2025)
by: Zheng, Mingqian, et al.
Published: (2025)
Towards properly implementing Theory of Mind in AI systems: An account of four misconceptions
by: van der Meulen, Ramira, et al.
Published: (2025)
by: van der Meulen, Ramira, et al.
Published: (2025)
Beyond Theory of Mind in Robotics
by: Jung, Malte F.
Published: (2026)
by: Jung, Malte F.
Published: (2026)
Privy: Envisioning and Mitigating Privacy Risks for Consumer-facing AI Product Concepts
by: Lee, Hao-Ping, et al.
Published: (2025)
by: Lee, Hao-Ping, et al.
Published: (2025)
Spontaneous Theory of Mind for Artificial Intelligence
by: Gurney, Nikolos, et al.
Published: (2024)
by: Gurney, Nikolos, et al.
Published: (2024)
Using Learning Theories to Evolve Human-Centered XAI: Future Perspectives and Challenges
by: Cortinas-Lorenzo, Karina, et al.
Published: (2026)
by: Cortinas-Lorenzo, Karina, et al.
Published: (2026)
Un-Straightening Generative AI: How Queer Artists Surface and Challenge the Normativity of Generative AI Models
by: Taylor, Jordan, et al.
Published: (2025)
by: Taylor, Jordan, et al.
Published: (2025)
A Survey on Large Language Model Hallucination via a Creativity Perspective
by: Jiang, Xuhui, et al.
Published: (2024)
by: Jiang, Xuhui, et al.
Published: (2024)
VenusBench-Mobile: A Challenging and User-Centric Benchmark for Mobile GUI Agents with Capability Diagnostics
by: Gong, Yichen, et al.
Published: (2026)
by: Gong, Yichen, et al.
Published: (2026)
MLLM as a UI Judge: Benchmarking Multimodal LLMs for Predicting Human Perception of User Interfaces
by: Luera, Reuben A., et al.
Published: (2025)
by: Luera, Reuben A., et al.
Published: (2025)
Beyond Chat: a Framework for LLMs as Human-Centered Support Systems
by: Zhou, Zhiyin
Published: (2025)
by: Zhou, Zhiyin
Published: (2025)
Mind the Value-Action Gap: Do LLMs Act in Alignment with Their Values?
by: Shen, Hua, et al.
Published: (2025)
by: Shen, Hua, et al.
Published: (2025)
Human-Centered LLM-Agent User Interface: A Position Paper
by: Chin, Daniel, et al.
Published: (2024)
by: Chin, Daniel, et al.
Published: (2024)
DiagLink: A Dual-User Diagnostic Assistance System by Synergizing Experts with LLMs and Knowledge Graphs
by: Zhou, Zihan, et al.
Published: (2026)
by: Zhou, Zihan, et al.
Published: (2026)
Towards Human-Centered RegTech: Unpacking Professionals' Strategies and Needs for Using LLMs Safely
by: Hu, Siying, et al.
Published: (2025)
by: Hu, Siying, et al.
Published: (2025)
When Researchers Say Mental Model/Theory of Mind of AI, What Are They Really Talking About?
by: Yin, Xiaoyun, et al.
Published: (2025)
by: Yin, Xiaoyun, et al.
Published: (2025)
Effects of Theory of Mind and Prosocial Beliefs on Steering Human-Aligned Behaviors of LLMs in Ultimatum Games
by: Yadav, Neemesh, et al.
Published: (2025)
by: Yadav, Neemesh, et al.
Published: (2025)
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind
by: Lupu, Andrei, et al.
Published: (2025)
by: Lupu, Andrei, et al.
Published: (2025)
CHBench: A Cognitive Hierarchy Benchmark for Evaluating Strategic Reasoning Capability of LLMs
by: Liu, Hongtao, et al.
Published: (2025)
by: Liu, Hongtao, et al.
Published: (2025)
Rethinking XAI Evaluation: A Human-Centered Audit of Shapley Benchmarks in High-Stakes Settings
by: Silva, Inês Oliveira e, et al.
Published: (2026)
by: Silva, Inês Oliveira e, et al.
Published: (2026)
RealUserSim: Bridging the Reality Gap in Agent Benchmarking via Grounded User Simulation
by: Zhu, Ming, et al.
Published: (2026)
by: Zhu, Ming, et al.
Published: (2026)
Human-Centered Human-AI Interaction (HC-HAII): A Human-Centered AI Perspective
by: Xu, Wei
Published: (2025)
by: Xu, Wei
Published: (2025)
Towards Real-World Validity in Generative AI Benchmarks: Understanding and Designing Domain-Centered Evaluations for Journalism Practitioners
by: Li, Charlotte, et al.
Published: (2025)
by: Li, Charlotte, et al.
Published: (2025)
Large Language Models in Peer-Run Community Behavioral Health Services: Understanding Peer Specialists and Service Users' Perspectives on Opportunities, Risks, and Mitigation Strategies
by: Peng, Cindy, et al.
Published: (2026)
by: Peng, Cindy, et al.
Published: (2026)
MobileAgentBench: An Efficient and User-Friendly Benchmark for Mobile LLM Agents
by: Wang, Luyuan, et al.
Published: (2024)
by: Wang, Luyuan, et al.
Published: (2024)
Exploring Big Five Personality and AI Capability Effects in LLM-Simulated Negotiation Dialogues
by: Cohen, Myke C., et al.
Published: (2025)
by: Cohen, Myke C., et al.
Published: (2025)
The Algorithmic Gaze of Image Quality Assessment: An Audit and Trace Ethnography of the LAION-Aesthetics Predictor
by: Taylor, Jordan, et al.
Published: (2026)
by: Taylor, Jordan, et al.
Published: (2026)
Heart2Mind: Human-Centered Contestable Psychiatric Disorder Diagnosis System using Wearable ECG Monitors
by: Nguyen, Hung, et al.
Published: (2025)
by: Nguyen, Hung, et al.
Published: (2025)
Situation Graph Prediction: Structured Perspective Inference for User Modeling
by: Shin, Jisung, et al.
Published: (2026)
by: Shin, Jisung, et al.
Published: (2026)
Similar Items
-
Situated, Dynamic, and Subjective: Envisioning the Design of Theory-of-Mind-Enabled Everyday AI with Industry Practitioners
by: Wang, Qiaosi, et al.
Published: (2026) -
Exploring the Innovation Opportunities for Pre-trained Models
by: Park, Minjung, et al.
Published: (2025) -
Minion: A Technology Probe to Explore How Users Negotiate Harmful Value Conflicts with AI Companions
by: Fan, Xianzhe, et al.
Published: (2024) -
Mutual Theory of Mind for Human-AI Communication
by: Wang, Qiaosi, et al.
Published: (2022) -
Promoting Critical Thinking With Domain-Specific Generative AI Provocations
by: von Davier, Thomas Şerban, et al.
Published: (2026)