ChatBench: From Static Benchmarks to Human-AI Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Chang, Serina, Anderson, Ashton, Hofman, Jake M. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When AI Gives Advice: Evaluating AI and Human Responses to Online Advice-Seeking for Well-Being
by: Kumar, Harsh, et al.
Published: (2025)
by: Kumar, Harsh, et al.
Published: (2025)
HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants
by: Sturgeon, Benjamin, et al.
Published: (2025)
by: Sturgeon, Benjamin, et al.
Published: (2025)
InvisibleBench: A Deployment Gate for Caregiving Relationship AI
by: Madad, Ali
Published: (2025)
by: Madad, Ali
Published: (2025)
Wayfinding through the AI wilderness: Mapping rhetorics of ChatGPT prompt writing on X (formerly Twitter) to promote critical AI literacies
by: Gupta, Anuj, et al.
Published: (2025)
by: Gupta, Anuj, et al.
Published: (2025)
Imperfectly Cooperative Human-AI Interactions: Comparing the Impacts of Human and AI Attributes in Simulated and User Studies
by: Cohen, Myke C., et al.
Published: (2026)
by: Cohen, Myke C., et al.
Published: (2026)
Human Decision-making is Susceptible to AI-driven Manipulation
by: Sabour, Sahand, et al.
Published: (2025)
by: Sabour, Sahand, et al.
Published: (2025)
How AI Ideas Affect the Creativity, Diversity, and Evolution of Human Ideas: Evidence From a Large, Dynamic Experiment
by: Ashkinaze, Joshua, et al.
Published: (2024)
by: Ashkinaze, Joshua, et al.
Published: (2024)
OpenTuringBench: An Open-Model-based Benchmark and Framework for Machine-Generated Text Detection and Attribution
by: La Cava, Lucio, et al.
Published: (2025)
by: La Cava, Lucio, et al.
Published: (2025)
TUX: Measuring Human--AI Tacit Understanding
by: Li, Yueshen, et al.
Published: (2026)
by: Li, Yueshen, et al.
Published: (2026)
From Reddit to Generative AI: Evaluating Large Language Models for Anxiety Support Fine-tuned on Social Media Data
by: Kursuncu, Ugur, et al.
Published: (2025)
by: Kursuncu, Ugur, et al.
Published: (2025)
Enhancing Depression-Diagnosis-Oriented Chat with Psychological State Tracking
by: Gu, Yiyang, et al.
Published: (2024)
by: Gu, Yiyang, et al.
Published: (2024)
An Investigation of Warning Erroneous Chat Translations in Cross-lingual Communication
by: Li, Yunmeng, et al.
Published: (2024)
by: Li, Yunmeng, et al.
Published: (2024)
Human-Centric NLP or AI-Centric Illusion?: A Critical Investigation
by: Spencer, Piyapath T
Published: (2024)
by: Spencer, Piyapath T
Published: (2024)
ChatGPT and U(X): A Rapid Review on Measuring the User Experience
by: Seaborn, Katie
Published: (2025)
by: Seaborn, Katie
Published: (2025)
To what extent is ChatGPT useful for language teacher lesson plan creation?
by: Dornburg, Alex, et al.
Published: (2024)
by: Dornburg, Alex, et al.
Published: (2024)
Conversational DNA: A New Visual Language for Understanding Dialogue Structure in Human and AI
by: Lin, Baihan
Published: (2025)
by: Lin, Baihan
Published: (2025)
Who's Asking? Simulating Role-Based Questions for Conversational AI Evaluation
by: Kaur, Navreet, et al.
Published: (2025)
by: Kaur, Navreet, et al.
Published: (2025)
GenAI Against Humanity: Nefarious Applications of Generative Artificial Intelligence and Large Language Models
by: Ferrara, Emilio
Published: (2023)
by: Ferrara, Emilio
Published: (2023)
Leading Across the Spectrum of Human-AI Relationships: A Conceptual Framework for Increasingly Heterogeneous Teams
by: Jadad, Alejandro R.
Published: (2026)
by: Jadad, Alejandro R.
Published: (2026)
From Google Gemini to OpenAI Q* (Q-Star): A Survey of Reshaping the Generative Artificial Intelligence (AI) Research Landscape
by: McIntosh, Timothy R., et al.
Published: (2023)
by: McIntosh, Timothy R., et al.
Published: (2023)
MentalChat16K: A Benchmark Dataset for Conversational Mental Health Assistance
by: Xu, Jia, et al.
Published: (2025)
by: Xu, Jia, et al.
Published: (2025)
From tools to thieves: Measuring and understanding public perceptions of AI through crowdsourced metaphors
by: Cheng, Myra, et al.
Published: (2025)
by: Cheng, Myra, et al.
Published: (2025)
If Eleanor Rigby Had Met ChatGPT: A Study on Loneliness in a Post-LLM World
by: de Wynter, Adrian
Published: (2024)
by: de Wynter, Adrian
Published: (2024)
Who Would Chatbots Vote For? Political Preferences of ChatGPT and Gemini in the 2024 European Union Elections
by: Haman, Michael, et al.
Published: (2024)
by: Haman, Michael, et al.
Published: (2024)
Red Teaming LLMs as Socio-Technical Practice: From Exploration and Data Creation to Evaluation
by: Garcia, Adriana Alvarado, et al.
Published: (2026)
by: Garcia, Adriana Alvarado, et al.
Published: (2026)
Human-Centered AI in Multidisciplinary Medical Discussions: Evaluating the Feasibility of a Chat-Based Approach to Case Assessment
by: Sawano, Shinnosuke, et al.
Published: (2025)
by: Sawano, Shinnosuke, et al.
Published: (2025)
Literary Narrative as Moral Probe : A Cross-System Framework for Evaluating AI Ethical Reasoning and Refusal Behavior
by: Flynn, David C.
Published: (2026)
by: Flynn, David C.
Published: (2026)
From Divergence to Consensus: Evaluating the Role of Large Language Models in Facilitating Agreement through Adaptive Strategies
by: Triantafyllopoulos, Loukas, et al.
Published: (2025)
by: Triantafyllopoulos, Loukas, et al.
Published: (2025)
The Levers of Political Persuasion with Conversational AI
by: Hackenburg, Kobi, et al.
Published: (2025)
by: Hackenburg, Kobi, et al.
Published: (2025)
Bottom-Up Perspectives on AI Governance: Insights from User Reviews of AI Products
by: Pasch, Stefan
Published: (2025)
by: Pasch, Stefan
Published: (2025)
Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence
by: McIntosh, Timothy R., et al.
Published: (2024)
by: McIntosh, Timothy R., et al.
Published: (2024)
Human Preferences for Constructive Interactions in Language Model Alignment
by: Kyrychenko, Yara, et al.
Published: (2025)
by: Kyrychenko, Yara, et al.
Published: (2025)
Lessons From an App Update at Replika AI: Identity Discontinuity in Human-AI Relationships
by: De Freitas, Julian, et al.
Published: (2024)
by: De Freitas, Julian, et al.
Published: (2024)
Human-Centred LLM Privacy Audits: Findings and Frictions
by: Staufer, Dimitri, et al.
Published: (2026)
by: Staufer, Dimitri, et al.
Published: (2026)
Generative AI Perceptions: A Survey to Measure the Perceptions of Faculty, Staff, and Students on Generative AI Tools in Academia
by: Amani, Sara, et al.
Published: (2023)
by: Amani, Sara, et al.
Published: (2023)
MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes
by: Chiu, Yu Ying, et al.
Published: (2025)
by: Chiu, Yu Ying, et al.
Published: (2025)
Mind the Style: Impact of Communication Style on Human-Chatbot Interaction
by: Derner, Erik, et al.
Published: (2026)
by: Derner, Erik, et al.
Published: (2026)
Does AI Coaching Prepare us for Workplace Negotiations?
by: Duddu, Veda, et al.
Published: (2025)
by: Duddu, Veda, et al.
Published: (2025)
AI-Generated Slides: Are They Good? Can Students Tell?
by: Leinonen, Juho, et al.
Published: (2026)
by: Leinonen, Juho, et al.
Published: (2026)
Designing KRIYA: An AI Companion for Wellbeing Self-Reflection
by: Zhu, Shanshan, et al.
Published: (2026)
by: Zhu, Shanshan, et al.
Published: (2026)
Similar Items
-
When AI Gives Advice: Evaluating AI and Human Responses to Online Advice-Seeking for Well-Being
by: Kumar, Harsh, et al.
Published: (2025) -
HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants
by: Sturgeon, Benjamin, et al.
Published: (2025) -
InvisibleBench: A Deployment Gate for Caregiving Relationship AI
by: Madad, Ali
Published: (2025) -
Wayfinding through the AI wilderness: Mapping rhetorics of ChatGPT prompt writing on X (formerly Twitter) to promote critical AI literacies
by: Gupta, Anuj, et al.
Published: (2025) -
Imperfectly Cooperative Human-AI Interactions: Comparing the Impacts of Human and AI Attributes in Simulated and User Studies
by: Cohen, Myke C., et al.
Published: (2026)