Challenges in Trustworthy Human Evaluation of Chatbots
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Wenting, Rush, Alexander M., Goyal, Tanya |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Framework for Evaluating Appropriateness, Trustworthiness, and Safety in Mental Wellness AI Chatbots
by: Chen, Lucia, et al.
Published: (2024)
by: Chen, Lucia, et al.
Published: (2024)
How Do Teachers Create Pedagogical Chatbots?: Current Practices and Challenges
by: Yoo, Minju, et al.
Published: (2025)
by: Yoo, Minju, et al.
Published: (2025)
A Checklist for Trustworthy, Safe, and User-Friendly Mental Health Chatbots
by: Haran, Shreya, et al.
Published: (2026)
by: Haran, Shreya, et al.
Published: (2026)
Revisiting Human Information Foraging: Adaptations for LLM-based Chatbots
by: Ragavan, Sruti Srinivasa, et al.
Published: (2024)
by: Ragavan, Sruti Srinivasa, et al.
Published: (2024)
Exploring the Effects of Chatbot Anthropomorphism and Human Empathy on Human Prosocial Behavior Toward Chatbots
by: Li, Jingshu, et al.
Published: (2025)
by: Li, Jingshu, et al.
Published: (2025)
Structure Matters: Evaluating Multi-Agents Orchestration in Generative Therapeutic Chatbots
by: Elahimanesh, Sina, et al.
Published: (2026)
by: Elahimanesh, Sina, et al.
Published: (2026)
AI Chatbots or Human Therapists? Belief-Based Predictors of Mental Health Help-Seeking Intentions in the Age of Generative AI
by: Park, Junsang, et al.
Published: (2025)
by: Park, Junsang, et al.
Published: (2025)
A Scoping Study of Evaluation Practices for Responsible AI Tools: Steps Towards Effectiveness Evaluations
by: Berman, Glen, et al.
Published: (2024)
by: Berman, Glen, et al.
Published: (2024)
Designing an LLM-Based Behavioral Activation Chatbot for Young People with Depression: Insights from an Evaluation with Artificial Users and Clinical Experts
by: Kuhlmeier, Florian Onur, et al.
Published: (2025)
by: Kuhlmeier, Florian Onur, et al.
Published: (2025)
Evaluating an LLM-Powered Chatbot for Cognitive Restructuring: Insights from Mental Health Professionals
by: Wang, Yinzhou, et al.
Published: (2025)
by: Wang, Yinzhou, et al.
Published: (2025)
Chatbot apologies: Beyond bullshit
by: Magnus, P. D., et al.
Published: (2025)
by: Magnus, P. D., et al.
Published: (2025)
Engineering Trustworthy Automation: Design Principles and Evaluation for AutoML Tools for Novices
by: Thys, Jarne, et al.
Published: (2025)
by: Thys, Jarne, et al.
Published: (2025)
Building Trustworthy Cognitive Monitoring for Safety-Critical Human Tasks: A Phased Methodological Approach
by: Grzeszczuk, Maciej, et al.
Published: (2025)
by: Grzeszczuk, Maciej, et al.
Published: (2025)
A Mixed-Methods Evaluation of LLM-Based Chatbots for Menopause
by: Deva, Roshini, et al.
Published: (2025)
by: Deva, Roshini, et al.
Published: (2025)
Is AI mingling or bullying me? Exploring User Interactions with a Chatbot in China
by: Chen, Nuo, et al.
Published: (2025)
by: Chen, Nuo, et al.
Published: (2025)
Exploring Socio-Cultural Challenges and Opportunities in Designing Mental Health Chatbots for Adolescents in India
by: Sehgal, Neil K. R., et al.
Published: (2025)
by: Sehgal, Neil K. R., et al.
Published: (2025)
Chatbot Companionship: A Mixed-Methods Study of Companion Chatbot Usage Patterns and Their Relationship to Loneliness in Active Users
by: Liu, Auren R., et al.
Published: (2024)
by: Liu, Auren R., et al.
Published: (2024)
The Dynamics of Delusion: Modeling Bidirectional False Belief Amplification in Human-Chatbot Dialogue
by: Mehta, Ashish, et al.
Published: (2026)
by: Mehta, Ashish, et al.
Published: (2026)
Evaluating the Experience of LGBTQ+ People Using Large Language Model Based Chatbots for Mental Health Support
by: Ma, Zilin, et al.
Published: (2024)
by: Ma, Zilin, et al.
Published: (2024)
Evaluating Trust in AI, Human, and Co-produced Feedback Among Undergraduate Students
by: Zhang, Audrey, et al.
Published: (2025)
by: Zhang, Audrey, et al.
Published: (2025)
Bridging the Cognitive Gap: Co-Designing and Evaluating a Voice-Enabled Community Chatbot for Older Adults
by: Chen, Feng, et al.
Published: (2026)
by: Chen, Feng, et al.
Published: (2026)
Improving Public Service Chatbot Design and Civic Impact: Investigation of Citizens' Perceptions of a Metro City 311 Chatbot
by: Zhou, Jieyu, et al.
Published: (2025)
by: Zhou, Jieyu, et al.
Published: (2025)
Creating, Using and Assessing a Generative-AI-Based Human-Chatbot-Dialogue Dataset with User-Interaction Learning Capabilities
by: Cuzzocrea, Alfredo, et al.
Published: (2025)
by: Cuzzocrea, Alfredo, et al.
Published: (2025)
How AI and Human Behaviors Shape Psychosocial Effects of Extended Chatbot Use: A Longitudinal Randomized Controlled Study
by: Fang, Cathy Mengying, et al.
Published: (2025)
by: Fang, Cathy Mengying, et al.
Published: (2025)
AniBalloons: Animated Chat Balloons as Affective Augmentation for Social Messaging and Chatbot Interaction
by: An, Pengcheng, et al.
Published: (2024)
by: An, Pengcheng, et al.
Published: (2024)
Facilitating Asynchronous Idea Generation and Selection with Chatbots
by: Shin, Joongi, et al.
Published: (2025)
by: Shin, Joongi, et al.
Published: (2025)
Toward a Human-Centered Evaluation Framework for Trustworthy LLM-Powered GUI Agents
by: Chen, Chaoran, et al.
Published: (2025)
by: Chen, Chaoran, et al.
Published: (2025)
Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots
by: Zou, Huiqi, et al.
Published: (2024)
by: Zou, Huiqi, et al.
Published: (2024)
Supporting Effective Goal Setting with LLM-Based Chatbots
by: Schimpf, Michel, et al.
Published: (2026)
by: Schimpf, Michel, et al.
Published: (2026)
Disclose with Care: Designing Privacy Controls in Interview Chatbots
by: Li, Ziwen, et al.
Published: (2026)
by: Li, Ziwen, et al.
Published: (2026)
Users' Mental Models of Generative AI Chatbot Ecosystems
by: Wang, Xingyi, et al.
Published: (2025)
by: Wang, Xingyi, et al.
Published: (2025)
Exploring Sidewalk Sheds in New York City through Chatbot Surveys and Human Computer Interaction
by: Li, Junyi, et al.
Published: (2026)
by: Li, Junyi, et al.
Published: (2026)
A Piece of Theatre: Investigating How Teachers Design LLM Chatbots to Assist Adolescent Cyberbullying Education
by: Hedderich, Michael A., et al.
Published: (2024)
by: Hedderich, Michael A., et al.
Published: (2024)
Private Yet Social: How LLM Chatbots Support and Challenge Eating Disorder Recovery
by: Choi, Ryuhaerang, et al.
Published: (2024)
by: Choi, Ryuhaerang, et al.
Published: (2024)
Enhancing Critical Thinking in Education by means of a Socratic Chatbot
by: Favero, Lucile, et al.
Published: (2024)
by: Favero, Lucile, et al.
Published: (2024)
Watching AI Think: User Perceptions of Visible Thinking in Chatbots
by: Cox, Samuel Rhys, et al.
Published: (2026)
by: Cox, Samuel Rhys, et al.
Published: (2026)
Generating HomeAssistant Automations Using an LLM-based Chatbot
by: Giudici, Mathyas, et al.
Published: (2025)
by: Giudici, Mathyas, et al.
Published: (2025)
Chatbots language design: the influence of language variation on user experience
by: Chaves, Ana Paula, et al.
Published: (2021)
by: Chaves, Ana Paula, et al.
Published: (2021)
Do We Know What They Know We Know? Calibrating Student Trust in AI and Human Responses Through Mutual Theory of Mind
by: Pal, Olivia, et al.
Published: (2026)
by: Pal, Olivia, et al.
Published: (2026)
Chatbots for Data Collection in Surveys: A Comparison of Four Theory-Based Interview Probes
by: Jacobsen, Rune M., et al.
Published: (2025)
by: Jacobsen, Rune M., et al.
Published: (2025)
Similar Items
-
A Framework for Evaluating Appropriateness, Trustworthiness, and Safety in Mental Wellness AI Chatbots
by: Chen, Lucia, et al.
Published: (2024) -
How Do Teachers Create Pedagogical Chatbots?: Current Practices and Challenges
by: Yoo, Minju, et al.
Published: (2025) -
A Checklist for Trustworthy, Safe, and User-Friendly Mental Health Chatbots
by: Haran, Shreya, et al.
Published: (2026) -
Revisiting Human Information Foraging: Adaptations for LLM-based Chatbots
by: Ragavan, Sruti Srinivasa, et al.
Published: (2024) -
Exploring the Effects of Chatbot Anthropomorphism and Human Empathy on Human Prosocial Behavior Toward Chatbots
by: Li, Jingshu, et al.
Published: (2025)