Saved in:
| Main Authors: | Tenzer, Helene, Abidi, Oumnia, Feuerriegel, Stefan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2509.11921 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
NARRA-Gym for Evaluating Interactive Narrative Agents
by: Huang, Yue, et al.
Published: (2026)
by: Huang, Yue, et al.
Published: (2026)
Assessing LLMs in Art Contexts: Critique Generation and Theory of Mind Evaluation
by: Arita, Takaya, et al.
Published: (2025)
by: Arita, Takaya, et al.
Published: (2025)
Your Students Don't Use LLMs Like You Wish They Did
by: Kobler, Sebastian, et al.
Published: (2026)
by: Kobler, Sebastian, et al.
Published: (2026)
Real or Robotic? Assessing Whether LLMs Accurately Simulate Qualities of Human Responses in Dialogue
by: Ivey, Jonathan, et al.
Published: (2024)
by: Ivey, Jonathan, et al.
Published: (2024)
When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation
by: Badawi, Abeer, et al.
Published: (2025)
by: Badawi, Abeer, et al.
Published: (2025)
Designing Computational Tools for Exploring Causal Relationships in Qualitative Data
by: Meng, Han, et al.
Published: (2026)
by: Meng, Han, et al.
Published: (2026)
"Ownership, Not Just Happy Talk": Co-Designing a Participatory Large Language Model for Journalism
by: Tseng, Emily, et al.
Published: (2025)
by: Tseng, Emily, et al.
Published: (2025)
Human-centered NLP Fact-checking: Co-Designing with Fact-checkers using Matchmaking for AI
by: Liu, Houjiang, et al.
Published: (2023)
by: Liu, Houjiang, et al.
Published: (2023)
Design and consensus content validity of the questionnaire for b-learning education: A 2-Tuple Fuzzy Linguistic Delphi based Decision Support Tool
by: Montes, Rosana, et al.
Published: (2024)
by: Montes, Rosana, et al.
Published: (2024)
Bottom-Up Perspectives on AI Governance: Insights from User Reviews of AI Products
by: Pasch, Stefan
Published: (2025)
by: Pasch, Stefan
Published: (2025)
LLM Content Moderation and User Satisfaction: Evidence from Response Refusals in Chatbot Arena
by: Pasch, Stefan
Published: (2025)
by: Pasch, Stefan
Published: (2025)
Evidence of conceptual mastery in the application of rules by Large Language Models
by: Nunes, José Luiz, et al.
Published: (2025)
by: Nunes, José Luiz, et al.
Published: (2025)
Can LLMs Reason About Trust?: A Pilot Study
by: Debnath, Anushka, et al.
Published: (2025)
by: Debnath, Anushka, et al.
Published: (2025)
First Contact with Dark Patterns and Deceptive Designs in Chinese and Japanese Free-to-Play Mobile Games
by: Zhang, Gloria Xiaodan, et al.
Published: (2025)
by: Zhang, Gloria Xiaodan, et al.
Published: (2025)
Evidence of a log scaling law for political persuasion with large language models
by: Hackenburg, Kobi, et al.
Published: (2024)
by: Hackenburg, Kobi, et al.
Published: (2024)
The Reliability of LLMs for Medical Diagnosis: An Examination of Consistency, Manipulation, and Contextual Awareness
by: Subedi, Krishna
Published: (2025)
by: Subedi, Krishna
Published: (2025)
A Comprehensive Survey of Bias in LLMs: Current Landscape and Future Directions
by: Ranjan, Rajesh, et al.
Published: (2024)
by: Ranjan, Rajesh, et al.
Published: (2024)
Seeing Like an AI: How LLMs Apply (and Misapply) Wikipedia Neutrality Norms
by: Ashkinaze, Joshua, et al.
Published: (2024)
by: Ashkinaze, Joshua, et al.
Published: (2024)
Overreliance on AI in Information-seeking from Video Content
by: Møller, Anders Giovanni, et al.
Published: (2026)
by: Møller, Anders Giovanni, et al.
Published: (2026)
Humanizing LLMs: A Survey of Psychological Measurements with Tools, Datasets, and Human-Agent Applications
by: Dong, Wenhan, et al.
Published: (2025)
by: Dong, Wenhan, et al.
Published: (2025)
Red Teaming LLMs as Socio-Technical Practice: From Exploration and Data Creation to Evaluation
by: Garcia, Adriana Alvarado, et al.
Published: (2026)
by: Garcia, Adriana Alvarado, et al.
Published: (2026)
From "Help" to Helpful: A Hierarchical Assessment of LLMs in Mental e-Health Applications
by: Steigerwald, Philipp, et al.
Published: (2026)
by: Steigerwald, Philipp, et al.
Published: (2026)
How Persuasive Could LLMs Be? A First Study Combining Linguistic-Rhetorical Analysis and User Experiments
by: Raffini, Daniel, et al.
Published: (2025)
by: Raffini, Daniel, et al.
Published: (2025)
Understanding the LLM-ification of CHI: Unpacking the Impact of LLMs at CHI through a Systematic Literature Review
by: Pang, Rock Yuren, et al.
Published: (2025)
by: Pang, Rock Yuren, et al.
Published: (2025)
Designing KRIYA: An AI Companion for Wellbeing Self-Reflection
by: Zhu, Shanshan, et al.
Published: (2026)
by: Zhu, Shanshan, et al.
Published: (2026)
Toward Cultural Interpretability: A Linguistic Anthropological Framework for Describing and Evaluating Large Language Models (LLMs)
by: Jones, Graham M., et al.
Published: (2024)
by: Jones, Graham M., et al.
Published: (2024)
How AI Ideas Affect the Creativity, Diversity, and Evolution of Human Ideas: Evidence From a Large, Dynamic Experiment
by: Ashkinaze, Joshua, et al.
Published: (2024)
by: Ashkinaze, Joshua, et al.
Published: (2024)
What Do LLMs Associate with Your Name? A Human-Centered Black-Box Audit of Personal Data
by: Staufer, Dimitri, et al.
Published: (2026)
by: Staufer, Dimitri, et al.
Published: (2026)
Designing Explainable AI for Healthcare Reviews: Guidance on Adoption and Trust
by: Alamoudi, Eman, et al.
Published: (2026)
by: Alamoudi, Eman, et al.
Published: (2026)
Patterns vs. Patients: Evaluating LLMs against Mental Health Professionals on Personality Disorder Diagnosis through First-Person Narratives
by: Drożdż, Karolina, et al.
Published: (2025)
by: Drożdż, Karolina, et al.
Published: (2025)
The Hidden Language of Harm: Examining the Role of Emojis in Harmful Online Communication and Content Moderation
by: Zhou, Yuhang, et al.
Published: (2025)
by: Zhou, Yuhang, et al.
Published: (2025)
What is Stigma Attributed to? A Theory-Grounded, Expert-Annotated Interview Corpus for Demystifying Mental-Health Stigma
by: Meng, Han, et al.
Published: (2025)
by: Meng, Han, et al.
Published: (2025)
Epistemological Fault Lines Between Human and Artificial Intelligence
by: Quattrociocchi, Walter, et al.
Published: (2025)
by: Quattrociocchi, Walter, et al.
Published: (2025)
"Would You Want an AI Tutor?" Understanding Stakeholder Perceptions of LLM-based Systems in the Classroom
by: Fuligni, Caterina, et al.
Published: (2025)
by: Fuligni, Caterina, et al.
Published: (2025)
When Algorithms Meet Artists: Semantic Compression of Artists' Concerns in the Public AI-Art Debate
by: Mukherjee-Gandhi, Ariya, et al.
Published: (2025)
by: Mukherjee-Gandhi, Ariya, et al.
Published: (2025)
A Systematic Review on the Evaluation of Large Language Models in Theory of Mind Tasks
by: Sarıtaş, Karahan, et al.
Published: (2025)
by: Sarıtaş, Karahan, et al.
Published: (2025)
How do datasets, developers, and models affect biases in a low-resourced language?: The Case of the Bengali Language
by: Das, Dipto, et al.
Published: (2025)
by: Das, Dipto, et al.
Published: (2025)
Human Capital Visualization using Speech Amount during Meetings
by: Hashimoto, Ekai, et al.
Published: (2025)
by: Hashimoto, Ekai, et al.
Published: (2025)
Enhancing Mathematics Learning for Hard-of-Hearing Students Through Real-Time Palestinian Sign Language Recognition: A New Dataset
by: Khandaqji, Fidaa, et al.
Published: (2025)
by: Khandaqji, Fidaa, et al.
Published: (2025)
Longitudinal Monitoring of LLM Content Moderation of Social Issues
by: Dai, Yunlang, et al.
Published: (2025)
by: Dai, Yunlang, et al.
Published: (2025)
Similar Items
-
NARRA-Gym for Evaluating Interactive Narrative Agents
by: Huang, Yue, et al.
Published: (2026) -
Assessing LLMs in Art Contexts: Critique Generation and Theory of Mind Evaluation
by: Arita, Takaya, et al.
Published: (2025) -
Your Students Don't Use LLMs Like You Wish They Did
by: Kobler, Sebastian, et al.
Published: (2026) -
Real or Robotic? Assessing Whether LLMs Accurately Simulate Qualities of Human Responses in Dialogue
by: Ivey, Jonathan, et al.
Published: (2024) -
When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation
by: Badawi, Abeer, et al.
Published: (2025)