Towards Dynamic Theory of Mind: Evaluating LLM Adaptation to Temporal Evolution of Human States
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xiao, Yang, Wang, Jiashuo, Xu, Qiancheng, Song, Changhe, Xu, Chunpu, Cheng, Yi, Li, Wenjie, Liu, Pengfei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards a Client-Centered Assessment of LLM Therapists by Client Simulation
von: Wang, Jiashuo, et al.
Veröffentlicht: (2024)
von: Wang, Jiashuo, et al.
Veröffentlicht: (2024)
How Far Are LLMs from Believable AI? A Benchmark for Evaluating the Believability of Human Behavior Simulation
von: Xiao, Yang, et al.
Veröffentlicht: (2023)
von: Xiao, Yang, et al.
Veröffentlicht: (2023)
Enhancing User Engagement in Socially-Driven Dialogue through Interactive LLM Alignments
von: Wang, Jiashuo, et al.
Veröffentlicht: (2025)
von: Wang, Jiashuo, et al.
Veröffentlicht: (2025)
SCALE: Selective Resource Allocation for Overcoming Performance Bottlenecks in Mathematical Test-time Scaling
von: Xiao, Yang, et al.
Veröffentlicht: (2025)
von: Xiao, Yang, et al.
Veröffentlicht: (2025)
LIMOPro: Reasoning Refinement for Efficient and Effective Test-time Scaling
von: Xiao, Yang, et al.
Veröffentlicht: (2025)
von: Xiao, Yang, et al.
Veröffentlicht: (2025)
Mitigating Unhelpfulness in Emotional Support Conversations with Multifaceted AI Feedback
von: Wang, Jiashuo, et al.
Veröffentlicht: (2024)
von: Wang, Jiashuo, et al.
Veröffentlicht: (2024)
Evaluating Large Language Models in Theory of Mind Tasks
von: Kosinski, Michal
Veröffentlicht: (2023)
von: Kosinski, Michal
Veröffentlicht: (2023)
Assessing LLMs in Art Contexts: Critique Generation and Theory of Mind Evaluation
von: Arita, Takaya, et al.
Veröffentlicht: (2025)
von: Arita, Takaya, et al.
Veröffentlicht: (2025)
Exploring the Human-LLM Synergy in Advancing Theory-driven Qualitative Analysis
von: Meng, Han, et al.
Veröffentlicht: (2024)
von: Meng, Han, et al.
Veröffentlicht: (2024)
A Systematic Review on the Evaluation of Large Language Models in Theory of Mind Tasks
von: Sarıtaş, Karahan, et al.
Veröffentlicht: (2025)
von: Sarıtaş, Karahan, et al.
Veröffentlicht: (2025)
Mind the Gap: Pitfalls of LLM Alignment with Asian Public Opinion
von: Shankar, Hari, et al.
Veröffentlicht: (2026)
von: Shankar, Hari, et al.
Veröffentlicht: (2026)
JiraiBench: A Bilingual Benchmark for Evaluating Large Language Models' Detection of Human Self-Destructive Behavior Content in Jirai Community
von: Xiao, Yunze, et al.
Veröffentlicht: (2025)
von: Xiao, Yunze, et al.
Veröffentlicht: (2025)
From Individuals to Interactions: Benchmarking Gender Bias in Multimodal Large Language Models from the Lens of Social Relationship
von: Xu, Yue, et al.
Veröffentlicht: (2025)
von: Xu, Yue, et al.
Veröffentlicht: (2025)
Challenges and Innovations in LLM-Powered Fake News Detection: A Synthesis of Approaches and Future Directions
von: Yi, Jingyuan, et al.
Veröffentlicht: (2025)
von: Yi, Jingyuan, et al.
Veröffentlicht: (2025)
CARE: Causality Reasoning for Empathetic Responses by Conditional Graph Generation
von: Wang, Jiashuo, et al.
Veröffentlicht: (2022)
von: Wang, Jiashuo, et al.
Veröffentlicht: (2022)
PEToolLLM: Towards Personalized Tool Learning in Large Language Models
von: Xu, Qiancheng, et al.
Veröffentlicht: (2025)
von: Xu, Qiancheng, et al.
Veröffentlicht: (2025)
MoVa: Towards Generalizable Classification of Human Morals and Values
von: Chen, Ziyu, et al.
Veröffentlicht: (2025)
von: Chen, Ziyu, et al.
Veröffentlicht: (2025)
Are Vision Language Models Cross-Cultural Theory of Mind Reasoners?
von: Nazi, Zabir Al, et al.
Veröffentlicht: (2025)
von: Nazi, Zabir Al, et al.
Veröffentlicht: (2025)
RAGAT-Mind: A Multi-Granular Modeling Approach for Rumor Detection Based on MindSpore
von: Qin, Zhenkai, et al.
Veröffentlicht: (2025)
von: Qin, Zhenkai, et al.
Veröffentlicht: (2025)
Triangulating Temporal Dynamics in Multilingual Swiss Online News
von: Victor, Bros, et al.
Veröffentlicht: (2026)
von: Victor, Bros, et al.
Veröffentlicht: (2026)
Representation Learning to Study Temporal Dynamics in Tutorial Scaffolding
von: Borchers, Conrad, et al.
Veröffentlicht: (2026)
von: Borchers, Conrad, et al.
Veröffentlicht: (2026)
BeliefShift: Benchmarking Temporal Belief Consistency and Opinion Drift in LLM Agents
von: Myakala, Praveen Kumar, et al.
Veröffentlicht: (2026)
von: Myakala, Praveen Kumar, et al.
Veröffentlicht: (2026)
Tracking the Temporal Dynamics of News Coverage of Catastrophic and Violent Events
von: Lugos, Emily, et al.
Veröffentlicht: (2026)
von: Lugos, Emily, et al.
Veröffentlicht: (2026)
The Parrot Dilemma: Human-Labeled vs. LLM-augmented Data in Classification Tasks
von: Møller, Anders Giovanni, et al.
Veröffentlicht: (2023)
von: Møller, Anders Giovanni, et al.
Veröffentlicht: (2023)
Foresight Optimization for Strategic Reasoning in Large Language Models
von: Wang, Jiashuo, et al.
Veröffentlicht: (2026)
von: Wang, Jiashuo, et al.
Veröffentlicht: (2026)
Minding the Politeness Gap in Cross-cultural Communication
von: Machino, Yuka, et al.
Veröffentlicht: (2025)
von: Machino, Yuka, et al.
Veröffentlicht: (2025)
LLM or Human? Perceptions of Trust and Information Quality in Research Summaries
von: Akpinar, Nil-Jana, et al.
Veröffentlicht: (2026)
von: Akpinar, Nil-Jana, et al.
Veröffentlicht: (2026)
SHIELD: Evaluation and Defense Strategies for Copyright Compliance in LLM Text Generation
von: Liu, Xiaoze, et al.
Veröffentlicht: (2024)
von: Liu, Xiaoze, et al.
Veröffentlicht: (2024)
Human or LLM as Standardized Patients? A Comparative Study for Medical Education
von: Zhang, Bingquan, et al.
Veröffentlicht: (2025)
von: Zhang, Bingquan, et al.
Veröffentlicht: (2025)
Misalignment of LLM-Generated Personas with Human Perceptions in Low-Resource Settings
von: Prama, Tabia Tanzin, et al.
Veröffentlicht: (2025)
von: Prama, Tabia Tanzin, et al.
Veröffentlicht: (2025)
Agree to Disagree? A Meta-Evaluation of LLM Misgendering
von: Subramonian, Arjun, et al.
Veröffentlicht: (2025)
von: Subramonian, Arjun, et al.
Veröffentlicht: (2025)
Towards Implicit Bias Detection and Mitigation in Multi-Agent LLM Interactions
von: Borah, Angana, et al.
Veröffentlicht: (2024)
von: Borah, Angana, et al.
Veröffentlicht: (2024)
OATH-Frames: Characterizing Online Attitudes Towards Homelessness with LLM Assistants
von: Ranjit, Jaspreet, et al.
Veröffentlicht: (2024)
von: Ranjit, Jaspreet, et al.
Veröffentlicht: (2024)
Mind the Style: Impact of Communication Style on Human-Chatbot Interaction
von: Derner, Erik, et al.
Veröffentlicht: (2026)
von: Derner, Erik, et al.
Veröffentlicht: (2026)
LLMs are Biased Teachers: Evaluating LLM Bias in Personalized Education
von: Weissburg, Iain, et al.
Veröffentlicht: (2024)
von: Weissburg, Iain, et al.
Veröffentlicht: (2024)
Evaluating the Simulation of Human Personality-Driven Susceptibility to Misinformation with LLMs
von: Pratelli, Manuel, et al.
Veröffentlicht: (2025)
von: Pratelli, Manuel, et al.
Veröffentlicht: (2025)
PolicyLLM: Towards Excellent Comprehension of Public Policy for Large Language Models
von: Bao, Han, et al.
Veröffentlicht: (2026)
von: Bao, Han, et al.
Veröffentlicht: (2026)
Bias and Volatility: A Statistical Framework for Evaluating Large Language Model's Stereotypes and the Associated Generation Inconsistency
von: Liu, Yiran, et al.
Veröffentlicht: (2024)
von: Liu, Yiran, et al.
Veröffentlicht: (2024)
Probing Cultural Awareness in LLMs: A Case Study of Cross-Culture Aesthetic Stylistics
von: Wang, Jiashuo, et al.
Veröffentlicht: (2026)
von: Wang, Jiashuo, et al.
Veröffentlicht: (2026)
Few-shot Hate Speech Detection Based on the MindSpore Framework
von: Qin, Zhenkai, et al.
Veröffentlicht: (2025)
von: Qin, Zhenkai, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Towards a Client-Centered Assessment of LLM Therapists by Client Simulation
von: Wang, Jiashuo, et al.
Veröffentlicht: (2024) -
How Far Are LLMs from Believable AI? A Benchmark for Evaluating the Believability of Human Behavior Simulation
von: Xiao, Yang, et al.
Veröffentlicht: (2023) -
Enhancing User Engagement in Socially-Driven Dialogue through Interactive LLM Alignments
von: Wang, Jiashuo, et al.
Veröffentlicht: (2025) -
SCALE: Selective Resource Allocation for Overcoming Performance Bottlenecks in Mathematical Test-time Scaling
von: Xiao, Yang, et al.
Veröffentlicht: (2025) -
LIMOPro: Reasoning Refinement for Efficient and Effective Test-time Scaling
von: Xiao, Yang, et al.
Veröffentlicht: (2025)