Synthetic Heuristic Evaluation: A Comparison between AI- and Human-Powered Usability Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Zhong, Ruican, McDonald, David W., Hsieh, Gary |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Synthetic Cognitive Walkthrough: Aligning Large Language Model Performance with Human Cognitive Walkthrough
by: Zhong, Ruican, et al.
Published: (2025)
by: Zhong, Ruican, et al.
Published: (2025)
Controllable Complementarity: Subjective Preferences in Human-AI Collaboration
by: McDonald, Chase, et al.
Published: (2025)
by: McDonald, Chase, et al.
Published: (2025)
AI-Assisted Causal Pathway Diagram for Human-Centered Design
by: Zhong, Ruican, et al.
Published: (2024)
by: Zhong, Ruican, et al.
Published: (2024)
Levels of Autonomy for AI Agents
by: Feng, K. J. Kevin, et al.
Published: (2025)
by: Feng, K. J. Kevin, et al.
Published: (2025)
Is Seeing Believing? Evaluating Human Sensitivity to Synthetic Video
by: Wegmann, David, et al.
Published: (2026)
by: Wegmann, David, et al.
Published: (2026)
CoGrid & the Multi-User Gymnasium: A Framework for Multi-Agent Experimentation
by: McDonald, Chase, et al.
Published: (2026)
by: McDonald, Chase, et al.
Published: (2026)
Trusting Your AI Agent Emotionally and Cognitively: Development and Validation of a Semantic Differential Scale for AI Trust
by: Shang, Ruoxi, et al.
Published: (2024)
by: Shang, Ruoxi, et al.
Published: (2024)
Evaluating Usability and Engagement of Large Language Models in Virtual Reality for Traditional Scottish Curling
by: Lau, Ka Hei Carrie, et al.
Published: (2024)
by: Lau, Ka Hei Carrie, et al.
Published: (2024)
SPHERE: An Evaluation Card for Human-AI Systems
by: Ma, Qianou, et al.
Published: (2025)
by: Ma, Qianou, et al.
Published: (2025)
Generate, Evaluate, Iterate: Synthetic Data for Human-in-the-Loop Refinement of LLM Judges
by: Do, Hyo Jin, et al.
Published: (2025)
by: Do, Hyo Jin, et al.
Published: (2025)
A Risk Ontology for Evaluating AI-Powered Psychotherapy Virtual Agents
by: Steenstra, Ian, et al.
Published: (2025)
by: Steenstra, Ian, et al.
Published: (2025)
Evaluating Human-AI Collaboration: A Review and Methodological Framework
by: Fragiadakis, George, et al.
Published: (2024)
by: Fragiadakis, George, et al.
Published: (2024)
Evaluating Human-AI Interaction via Usability, User Experience and Acceptance Measures for MMM-C: A Creative AI System for Music Composition
by: Tchemeube, Renaud Bougueng, et al.
Published: (2025)
by: Tchemeube, Renaud Bougueng, et al.
Published: (2025)
An Approach to Grounding AI Model Evaluations in Human-derived Criteria
by: Mitts, Sasha
Published: (2025)
by: Mitts, Sasha
Published: (2025)
Interoceptive Divergence in Aesthetic Evaluation and Implications for Human-AI Alignment
by: Abe, Yoshia, et al.
Published: (2026)
by: Abe, Yoshia, et al.
Published: (2026)
Computer-Aided Tagging on Wikimedia Commons: Designing for Human-AI Collaboration in Open Knowledge Work
by: Yu, Yihan, et al.
Published: (2026)
by: Yu, Yihan, et al.
Published: (2026)
Investigating Multimodal Large Language Models to Support Usability Evaluation
by: Lubos, Sebastian, et al.
Published: (2025)
by: Lubos, Sebastian, et al.
Published: (2025)
Towards a Comprehensive Human-Centred Evaluation Framework for Explainable AI
by: Donoso-Guzmán, Ivania, et al.
Published: (2023)
by: Donoso-Guzmán, Ivania, et al.
Published: (2023)
How Human-Centered Explainable AI Interface Are Designed and Evaluated: A Systematic Survey
by: Nguyen, Thu, et al.
Published: (2024)
by: Nguyen, Thu, et al.
Published: (2024)
Towards Human-AI Deliberation: Design and Evaluation of LLM-Empowered Deliberative AI for AI-Assisted Decision-Making
by: Ma, Shuai, et al.
Published: (2024)
by: Ma, Shuai, et al.
Published: (2024)
Evaluating the Impact of AI-Powered Audiovisual Personalization on Learner Emotion, Focus, and Learning Outcomes
by: Wang, George Xi, et al.
Published: (2025)
by: Wang, George Xi, et al.
Published: (2025)
Synthetic Human Memories: AI-Edited Images and Videos Can Implant False Memories and Distort Recollection
by: Pataranutaporn, Pat, et al.
Published: (2024)
by: Pataranutaporn, Pat, et al.
Published: (2024)
Evaluating Trust in AI, Human, and Co-produced Feedback Among Undergraduate Students
by: Zhang, Audrey, et al.
Published: (2025)
by: Zhang, Audrey, et al.
Published: (2025)
Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing
by: Jafari, Kiana, et al.
Published: (2026)
by: Jafari, Kiana, et al.
Published: (2026)
Optimizing Generative AI's Accuracy and Transparency in Inductive Thematic Analysis: A Human-AI Comparison
by: Nyaaba, Matthew, et al.
Published: (2025)
by: Nyaaba, Matthew, et al.
Published: (2025)
Human-AI Interaction Alignment: Designing, Evaluating, and Evolving Value-Centered AI For Reciprocal Human-AI Futures
by: Shen, Hua, et al.
Published: (2025)
by: Shen, Hua, et al.
Published: (2025)
Evaluating Behavioral Alignment in Conflict Dialogue: A Multi-Dimensional Comparison of LLM Agents and Humans
by: Kwon, Deuksin, et al.
Published: (2025)
by: Kwon, Deuksin, et al.
Published: (2025)
On Evaluating Explanation Utility for Human-AI Decision Making in NLP
by: Chaleshtori, Fateme Hashemi, et al.
Published: (2024)
by: Chaleshtori, Fateme Hashemi, et al.
Published: (2024)
PaperTok: Exploring the Use of Generative AI for Creating Short-form Videos for Research Communication
by: Cristobal, Meziah Ruby, et al.
Published: (2026)
by: Cristobal, Meziah Ruby, et al.
Published: (2026)
Evaluating AI Alignment in LLMs: Output Analysis of Value Priorities Across 75 Models with Human Benchmarking
by: Lau, Gabriel Rongyang, et al.
Published: (2025)
by: Lau, Gabriel Rongyang, et al.
Published: (2025)
An Empirical Examination of the Evaluative AI Framework
by: Kornowicz, Jaroslaw
Published: (2024)
by: Kornowicz, Jaroslaw
Published: (2024)
Evaluating Human-AI Safety: A Framework for Measuring Harmful Capability Uplift
by: Vaccaro, Michelle, et al.
Published: (2026)
by: Vaccaro, Michelle, et al.
Published: (2026)
LLM-Powered Hierarchical Language Agent for Real-time Human-AI Coordination
by: Liu, Jijia, et al.
Published: (2023)
by: Liu, Jijia, et al.
Published: (2023)
Evaluation format, not model capability, drives triage failure in the assessment of consumer health AI
by: Navarro, David Fraile, et al.
Published: (2026)
by: Navarro, David Fraile, et al.
Published: (2026)
Evaluation of Architectural Synthesis Using Generative AI
by: Huang, Jingfei, et al.
Published: (2025)
by: Huang, Jingfei, et al.
Published: (2025)
Evaluating the Effects of AI Directors for Quest Selection
by: Yu, Kristen K., et al.
Published: (2024)
by: Yu, Kristen K., et al.
Published: (2024)
From Evidence to Decision: Exploring Evaluative AI
by: Le, Thao, et al.
Published: (2024)
by: Le, Thao, et al.
Published: (2024)
Beyond "Hallucinations": A Framework for Stable Human-AI Reasoning
by: Rosenbacke, Rikard, et al.
Published: (2025)
by: Rosenbacke, Rikard, et al.
Published: (2025)
Evaluating Human Trust in LLM-Based Planners: A Preliminary Study
by: Chen, Shenghui, et al.
Published: (2025)
by: Chen, Shenghui, et al.
Published: (2025)
Memory Power Asymmetry in Human-AI Relationships: Preserving Mutual Forgetting in the Digital Age
by: Dorri, Rasam, et al.
Published: (2025)
by: Dorri, Rasam, et al.
Published: (2025)
Similar Items
-
Synthetic Cognitive Walkthrough: Aligning Large Language Model Performance with Human Cognitive Walkthrough
by: Zhong, Ruican, et al.
Published: (2025) -
Controllable Complementarity: Subjective Preferences in Human-AI Collaboration
by: McDonald, Chase, et al.
Published: (2025) -
AI-Assisted Causal Pathway Diagram for Human-Centered Design
by: Zhong, Ruican, et al.
Published: (2024) -
Levels of Autonomy for AI Agents
by: Feng, K. J. Kevin, et al.
Published: (2025) -
Is Seeing Believing? Evaluating Human Sensitivity to Synthetic Video
by: Wegmann, David, et al.
Published: (2026)