Assessing LLMs in Art Contexts: Critique Generation and Theory of Mind Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Arita, Takaya, Zheng, Wenxian, Suzuki, Reiji, Akiba, Fuminori |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluating Large Language Models in Theory of Mind Tasks
by: Kosinski, Michal
Published: (2023)
by: Kosinski, Michal
Published: (2023)
A Systematic Review on the Evaluation of Large Language Models in Theory of Mind Tasks
by: Sarıtaş, Karahan, et al.
Published: (2025)
by: Sarıtaş, Karahan, et al.
Published: (2025)
Real or Robotic? Assessing Whether LLMs Accurately Simulate Qualities of Human Responses in Dialogue
by: Ivey, Jonathan, et al.
Published: (2024)
by: Ivey, Jonathan, et al.
Published: (2024)
When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation
by: Badawi, Abeer, et al.
Published: (2025)
by: Badawi, Abeer, et al.
Published: (2025)
When Algorithms Meet Artists: Semantic Compression of Artists' Concerns in the Public AI-Art Debate
by: Mukherjee-Gandhi, Ariya, et al.
Published: (2025)
by: Mukherjee-Gandhi, Ariya, et al.
Published: (2025)
An evolutionary model of personality traits related to cooperative behavior using a large language model
by: Suzuki, Reiji, et al.
Published: (2023)
by: Suzuki, Reiji, et al.
Published: (2023)
Decoding the Mind of Large Language Models: A Quantitative Evaluation of Ideology and Biases
by: Hirose, Manari, et al.
Published: (2025)
by: Hirose, Manari, et al.
Published: (2025)
Towards Better Health Conversations: The Benefits of Context-seeking
by: Sayres, Rory, et al.
Published: (2025)
by: Sayres, Rory, et al.
Published: (2025)
Designing LLMs for cultural sensitivity: Evidence from English-Japanese translation
by: Tenzer, Helene, et al.
Published: (2025)
by: Tenzer, Helene, et al.
Published: (2025)
Your Students Don't Use LLMs Like You Wish They Did
by: Kobler, Sebastian, et al.
Published: (2026)
by: Kobler, Sebastian, et al.
Published: (2026)
Exploring the Human-LLM Synergy in Advancing Theory-driven Qualitative Analysis
by: Meng, Han, et al.
Published: (2024)
by: Meng, Han, et al.
Published: (2024)
NARRA-Gym for Evaluating Interactive Narrative Agents
by: Huang, Yue, et al.
Published: (2026)
by: Huang, Yue, et al.
Published: (2026)
What is Stigma Attributed to? A Theory-Grounded, Expert-Annotated Interview Corpus for Demystifying Mental-Health Stigma
by: Meng, Han, et al.
Published: (2025)
by: Meng, Han, et al.
Published: (2025)
Mind the Style: Impact of Communication Style on Human-Chatbot Interaction
by: Derner, Erik, et al.
Published: (2026)
by: Derner, Erik, et al.
Published: (2026)
Multi-turn Evaluation of Anthropomorphic Behaviours in Large Language Models
by: Ibrahim, Lujain, et al.
Published: (2025)
by: Ibrahim, Lujain, et al.
Published: (2025)
Assessing LLM Response Quality in the Context of Technology-Facilitated Abuse
by: Prakash, Vijay, et al.
Published: (2026)
by: Prakash, Vijay, et al.
Published: (2026)
Theory of Mind and Self-Disclosure to CUIs
by: Cox, Samuel Rhys
Published: (2025)
by: Cox, Samuel Rhys
Published: (2025)
Red Teaming LLMs as Socio-Technical Practice: From Exploration and Data Creation to Evaluation
by: Garcia, Adriana Alvarado, et al.
Published: (2026)
by: Garcia, Adriana Alvarado, et al.
Published: (2026)
DaKultur: Evaluating the Cultural Awareness of Language Models for Danish with Native Speakers
by: Müller-Eberstein, Max, et al.
Published: (2025)
by: Müller-Eberstein, Max, et al.
Published: (2025)
SOTOPIA-$Ω$: Dynamic Strategy Injection Learning and Social Instruction Following Evaluation for Social Agents
by: Zhang, Wenyuan, et al.
Published: (2025)
by: Zhang, Wenyuan, et al.
Published: (2025)
Evaluating the capability of large language models to personalize science texts for diverse middle-school-age learners
by: Vaccaro Jr, Michael, et al.
Published: (2024)
by: Vaccaro Jr, Michael, et al.
Published: (2024)
Toward Cultural Interpretability: A Linguistic Anthropological Framework for Describing and Evaluating Large Language Models (LLMs)
by: Jones, Graham M., et al.
Published: (2024)
by: Jones, Graham M., et al.
Published: (2024)
Eskwai for Students: Generative AI Assistant for Legal Education in Ghana
by: Boateng, George, et al.
Published: (2026)
by: Boateng, George, et al.
Published: (2026)
Identity-related Speech Suppression in Generative AI Content Moderation
by: Proebsting, Grace, et al.
Published: (2024)
by: Proebsting, Grace, et al.
Published: (2024)
Patterns vs. Patients: Evaluating LLMs against Mental Health Professionals on Personality Disorder Diagnosis through First-Person Narratives
by: Drożdż, Karolina, et al.
Published: (2025)
by: Drożdż, Karolina, et al.
Published: (2025)
Kwame 2.0: Human-in-the-Loop Generative AI Teaching Assistant for Large Scale Online Coding Education in Africa
by: Boateng, George, et al.
Published: (2026)
by: Boateng, George, et al.
Published: (2026)
Open Models, Closed Minds? On Agents Capabilities in Mimicking Human Personalities through Open Large Language Models
by: La Cava, Lucio, et al.
Published: (2024)
by: La Cava, Lucio, et al.
Published: (2024)
Evaluating the Application of Large Language Models to Generate Feedback in Programming Education
by: Jacobs, Sven, et al.
Published: (2024)
by: Jacobs, Sven, et al.
Published: (2024)
Humanizing LLMs: A Survey of Psychological Measurements with Tools, Datasets, and Human-Agent Applications
by: Dong, Wenhan, et al.
Published: (2025)
by: Dong, Wenhan, et al.
Published: (2025)
Assessing the nature of large language models: A caution against anthropocentrism
by: Speed, Ann
Published: (2023)
by: Speed, Ann
Published: (2023)
Between Rules and Reality: On the Context Sensitivity of LLM Moral Judgment
by: Sauter, Adrian, et al.
Published: (2026)
by: Sauter, Adrian, et al.
Published: (2026)
An Investigation of Warning Erroneous Chat Translations in Cross-lingual Communication
by: Li, Yunmeng, et al.
Published: (2024)
by: Li, Yunmeng, et al.
Published: (2024)
Public Opinions About Copyright for AI-Generated Art: The Role of Egocentricity, Competition, and Experience
by: Lima, Gabriel, et al.
Published: (2024)
by: Lima, Gabriel, et al.
Published: (2024)
Can LLMs Reason About Trust?: A Pilot Study
by: Debnath, Anushka, et al.
Published: (2025)
by: Debnath, Anushka, et al.
Published: (2025)
From Reddit to Generative AI: Evaluating Large Language Models for Anxiety Support Fine-tuned on Social Media Data
by: Kursuncu, Ugur, et al.
Published: (2025)
by: Kursuncu, Ugur, et al.
Published: (2025)
The Reliability of LLMs for Medical Diagnosis: An Examination of Consistency, Manipulation, and Contextual Awareness
by: Subedi, Krishna
Published: (2025)
by: Subedi, Krishna
Published: (2025)
A Comprehensive Survey of Bias in LLMs: Current Landscape and Future Directions
by: Ranjan, Rajesh, et al.
Published: (2024)
by: Ranjan, Rajesh, et al.
Published: (2024)
Seeing Like an AI: How LLMs Apply (and Misapply) Wikipedia Neutrality Norms
by: Ashkinaze, Joshua, et al.
Published: (2024)
by: Ashkinaze, Joshua, et al.
Published: (2024)
From "Help" to Helpful: A Hierarchical Assessment of LLMs in Mental e-Health Applications
by: Steigerwald, Philipp, et al.
Published: (2026)
by: Steigerwald, Philipp, et al.
Published: (2026)
Exploring LLMs for Automated Generation and Adaptation of Questionnaires
by: Adhikari, Divya Mani, et al.
Published: (2025)
by: Adhikari, Divya Mani, et al.
Published: (2025)
Similar Items
-
Evaluating Large Language Models in Theory of Mind Tasks
by: Kosinski, Michal
Published: (2023) -
A Systematic Review on the Evaluation of Large Language Models in Theory of Mind Tasks
by: Sarıtaş, Karahan, et al.
Published: (2025) -
Real or Robotic? Assessing Whether LLMs Accurately Simulate Qualities of Human Responses in Dialogue
by: Ivey, Jonathan, et al.
Published: (2024) -
When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation
by: Badawi, Abeer, et al.
Published: (2025) -
When Algorithms Meet Artists: Semantic Compression of Artists' Concerns in the Public AI-Art Debate
by: Mukherjee-Gandhi, Ariya, et al.
Published: (2025)