A Systematic Review on the Evaluation of Large Language Models in Theory of Mind Tasks
Fuente:
arXiv
Salvato in:
| Autori principali: | Sarıtaş, Karahan, Tezören, Kıvanç, Durmazkeser, Yavuz |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Evaluating Large Language Models in Theory of Mind Tasks
di: Kosinski, Michal
Pubblicazione: (2023)
di: Kosinski, Michal
Pubblicazione: (2023)
Assessing LLMs in Art Contexts: Critique Generation and Theory of Mind Evaluation
di: Arita, Takaya, et al.
Pubblicazione: (2025)
di: Arita, Takaya, et al.
Pubblicazione: (2025)
Decoding the Mind of Large Language Models: A Quantitative Evaluation of Ideology and Biases
di: Hirose, Manari, et al.
Pubblicazione: (2025)
di: Hirose, Manari, et al.
Pubblicazione: (2025)
Multi-turn Evaluation of Anthropomorphic Behaviours in Large Language Models
di: Ibrahim, Lujain, et al.
Pubblicazione: (2025)
di: Ibrahim, Lujain, et al.
Pubblicazione: (2025)
Prompt Engineering Techniques for Mitigating Cultural Bias Against Arabs and Muslims in Large Language Models: A Systematic Review
di: Asseri, Bushra, et al.
Pubblicazione: (2025)
di: Asseri, Bushra, et al.
Pubblicazione: (2025)
Open Models, Closed Minds? On Agents Capabilities in Mimicking Human Personalities through Open Large Language Models
di: La Cava, Lucio, et al.
Pubblicazione: (2024)
di: La Cava, Lucio, et al.
Pubblicazione: (2024)
The Moral Machine Experiment on Large Language Models
di: Takemoto, Kazuhiro
Pubblicazione: (2023)
di: Takemoto, Kazuhiro
Pubblicazione: (2023)
The Perils & Promises of Fact-checking with Large Language Models
di: Quelle, Dorian, et al.
Pubblicazione: (2023)
di: Quelle, Dorian, et al.
Pubblicazione: (2023)
DaKultur: Evaluating the Cultural Awareness of Language Models for Danish with Native Speakers
di: Müller-Eberstein, Max, et al.
Pubblicazione: (2025)
di: Müller-Eberstein, Max, et al.
Pubblicazione: (2025)
Evaluating the Application of Large Language Models to Generate Feedback in Programming Education
di: Jacobs, Sven, et al.
Pubblicazione: (2024)
di: Jacobs, Sven, et al.
Pubblicazione: (2024)
"Ownership, Not Just Happy Talk": Co-Designing a Participatory Large Language Model for Journalism
di: Tseng, Emily, et al.
Pubblicazione: (2025)
di: Tseng, Emily, et al.
Pubblicazione: (2025)
ElectionSim: Massive Population Election Simulation Powered by Large Language Model Driven Agents
di: Zhang, Xinnong, et al.
Pubblicazione: (2024)
di: Zhang, Xinnong, et al.
Pubblicazione: (2024)
Toward Cultural Interpretability: A Linguistic Anthropological Framework for Describing and Evaluating Large Language Models (LLMs)
di: Jones, Graham M., et al.
Pubblicazione: (2024)
di: Jones, Graham M., et al.
Pubblicazione: (2024)
Evaluating Task-oriented Dialogue Systems: A Systematic Review of Measures, Constructs and their Operationalisations
di: Braggaar, Anouck, et al.
Pubblicazione: (2023)
di: Braggaar, Anouck, et al.
Pubblicazione: (2023)
From Divergence to Consensus: Evaluating the Role of Large Language Models in Facilitating Agreement through Adaptive Strategies
di: Triantafyllopoulos, Loukas, et al.
Pubblicazione: (2025)
di: Triantafyllopoulos, Loukas, et al.
Pubblicazione: (2025)
Language Models as Critical Thinking Tools: A Case Study of Philosophers
di: Ye, Andre, et al.
Pubblicazione: (2024)
di: Ye, Andre, et al.
Pubblicazione: (2024)
When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation
di: Badawi, Abeer, et al.
Pubblicazione: (2025)
di: Badawi, Abeer, et al.
Pubblicazione: (2025)
The Moral Gap of Large Language Models
di: Skorski, Maciej, et al.
Pubblicazione: (2025)
di: Skorski, Maciej, et al.
Pubblicazione: (2025)
Large Language Model Psychometrics: A Systematic Review of Evaluation, Validation, and Enhancement
di: Ye, Haoran, et al.
Pubblicazione: (2025)
di: Ye, Haoran, et al.
Pubblicazione: (2025)
Large Language Models as Psychological Simulators: A Methodological Guide
di: Lin, Zhicheng
Pubblicazione: (2025)
di: Lin, Zhicheng
Pubblicazione: (2025)
From Reddit to Generative AI: Evaluating Large Language Models for Anxiety Support Fine-tuned on Social Media Data
di: Kursuncu, Ugur, et al.
Pubblicazione: (2025)
di: Kursuncu, Ugur, et al.
Pubblicazione: (2025)
A Systematic Review on Prompt Engineering in Large Language Models for K-12 STEM Education
di: Chen, Eason, et al.
Pubblicazione: (2024)
di: Chen, Eason, et al.
Pubblicazione: (2024)
Exploring the Ethical Concerns in User Reviews of Mental Health Apps using Topic Modeling and Sentiment Analysis
di: Rahman, Mohammad Masudur, et al.
Pubblicazione: (2026)
di: Rahman, Mohammad Masudur, et al.
Pubblicazione: (2026)
Gender Trouble in Language Models: An Empirical Audit Guided by Gender Performativity Theory
di: Hafner, Franziska Sofia, et al.
Pubblicazione: (2025)
di: Hafner, Franziska Sofia, et al.
Pubblicazione: (2025)
More is More: Addition Bias in Large Language Models
di: Santagata, Luca, et al.
Pubblicazione: (2024)
di: Santagata, Luca, et al.
Pubblicazione: (2024)
Impacts of Anthropomorphizing Large Language Models in Learning Environments
di: Schaaff, Kristina, et al.
Pubblicazione: (2024)
di: Schaaff, Kristina, et al.
Pubblicazione: (2024)
What is Stigma Attributed to? A Theory-Grounded, Expert-Annotated Interview Corpus for Demystifying Mental-Health Stigma
di: Meng, Han, et al.
Pubblicazione: (2025)
di: Meng, Han, et al.
Pubblicazione: (2025)
Exploring the Human-LLM Synergy in Advancing Theory-driven Qualitative Analysis
di: Meng, Han, et al.
Pubblicazione: (2024)
di: Meng, Han, et al.
Pubblicazione: (2024)
Evidence of conceptual mastery in the application of rules by Large Language Models
di: Nunes, José Luiz, et al.
Pubblicazione: (2025)
di: Nunes, José Luiz, et al.
Pubblicazione: (2025)
On the Reliability of Large Language Models to Misinformed and Demographically-Informed Prompts
di: Aremu, Toluwani, et al.
Pubblicazione: (2024)
di: Aremu, Toluwani, et al.
Pubblicazione: (2024)
Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence
di: McIntosh, Timothy R., et al.
Pubblicazione: (2024)
di: McIntosh, Timothy R., et al.
Pubblicazione: (2024)
Empirical evidence of Large Language Model's influence on human spoken communication
di: Yakura, Hiromu, et al.
Pubblicazione: (2024)
di: Yakura, Hiromu, et al.
Pubblicazione: (2024)
Machine Learning Information Retrieval and Summarisation to Support Systematic Review on Outcomes Based Contracting
di: Bilal, Iman Munire, et al.
Pubblicazione: (2024)
di: Bilal, Iman Munire, et al.
Pubblicazione: (2024)
NARRA-Gym for Evaluating Interactive Narrative Agents
di: Huang, Yue, et al.
Pubblicazione: (2026)
di: Huang, Yue, et al.
Pubblicazione: (2026)
Implicit Personalization in Language Models: A Systematic Study
di: Jin, Zhijing, et al.
Pubblicazione: (2024)
di: Jin, Zhijing, et al.
Pubblicazione: (2024)
An Empirical Investigation of Gender Stereotype Representation in Large Language Models: The Italian Case
di: Giachino, Gioele, et al.
Pubblicazione: (2025)
di: Giachino, Gioele, et al.
Pubblicazione: (2025)
Getting in the Door: Streamlining Intake in Civil Legal Services with Large Language Models
di: Steenhuis, Quinten, et al.
Pubblicazione: (2024)
di: Steenhuis, Quinten, et al.
Pubblicazione: (2024)
Thinking with Many Minds: Using Large Language Models for Multi-Perspective Problem-Solving
di: Park, Sanghyun, et al.
Pubblicazione: (2025)
di: Park, Sanghyun, et al.
Pubblicazione: (2025)
Mind the Style: Impact of Communication Style on Human-Chatbot Interaction
di: Derner, Erik, et al.
Pubblicazione: (2026)
di: Derner, Erik, et al.
Pubblicazione: (2026)
Large-scale moral machine experiment on large language models
di: Ahmad, Muhammad Shahrul Zaim bin, et al.
Pubblicazione: (2024)
di: Ahmad, Muhammad Shahrul Zaim bin, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Evaluating Large Language Models in Theory of Mind Tasks
di: Kosinski, Michal
Pubblicazione: (2023) -
Assessing LLMs in Art Contexts: Critique Generation and Theory of Mind Evaluation
di: Arita, Takaya, et al.
Pubblicazione: (2025) -
Decoding the Mind of Large Language Models: A Quantitative Evaluation of Ideology and Biases
di: Hirose, Manari, et al.
Pubblicazione: (2025) -
Multi-turn Evaluation of Anthropomorphic Behaviours in Large Language Models
di: Ibrahim, Lujain, et al.
Pubblicazione: (2025) -
Prompt Engineering Techniques for Mitigating Cultural Bias Against Arabs and Muslims in Large Language Models: A Systematic Review
di: Asseri, Bushra, et al.
Pubblicazione: (2025)