A Scoping Study of Evaluation Practices for Responsible AI Tools: Steps Towards Effectiveness Evaluations
Fuente:
arXiv
Saved in:
| Main Authors: | Berman, Glen, Goyal, Nitesh, Madaio, Michael |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Troubling Taxonomies in GenAI Evaluation
by: Berman, Glen, et al.
Published: (2024)
by: Berman, Glen, et al.
Published: (2024)
The Case for "Thick Evaluations" of Cultural Representation in AI
by: Qadri, Rida, et al.
Published: (2025)
by: Qadri, Rida, et al.
Published: (2025)
"It Became My Buddy, But I'm Not Afraid to Disagree": A Multi-Session Study of UX Evaluators Collaborating with Conversational AI Assistants
by: Kuang, Emily, et al.
Published: (2026)
by: Kuang, Emily, et al.
Published: (2026)
Farsight: Fostering Responsible AI Awareness During AI Application Prototyping
by: Wang, Zijie J., et al.
Published: (2024)
by: Wang, Zijie J., et al.
Published: (2024)
AI Trust Reshaping Administrative Burdens: Understanding Trust-Burden Dynamics in LLM-Assisted Benefits Systems
by: Jo, Jeongwon, et al.
Published: (2025)
by: Jo, Jeongwon, et al.
Published: (2025)
Navigating Uncertainties: How GenAI Developers Document Their Models on Open-Source Platforms
by: Tang, Ningjing, et al.
Published: (2025)
by: Tang, Ningjing, et al.
Published: (2025)
"You have to prove the threat is real": Understanding the needs of Female Journalists and Activists to Document and Report Online Harassment
by: Goyal, Nitesh, et al.
Published: (2022)
by: Goyal, Nitesh, et al.
Published: (2022)
Implications of Regulations on the Use of AI and Generative AI for Human-Centered Responsible Artificial Intelligence
by: Constantinides, Marios, et al.
Published: (2024)
by: Constantinides, Marios, et al.
Published: (2024)
Challenges in Trustworthy Human Evaluation of Chatbots
by: Zhao, Wenting, et al.
Published: (2024)
by: Zhao, Wenting, et al.
Published: (2024)
Creating and Evaluating Personas Using Generative AI: A Scoping Review of 81 Articles
by: Amin, Danial, et al.
Published: (2025)
by: Amin, Danial, et al.
Published: (2025)
"Accessibility people, you go work on that thing of yours over there": Addressing Disability Inclusion in AI Product Organizations
by: Moharana, Sanika, et al.
Published: (2025)
by: Moharana, Sanika, et al.
Published: (2025)
Howzat? Appealing to Expert Judgement for Evaluating Human and AI Next-Step Hints for Novice Programmers
by: Brown, Neil C. C., et al.
Published: (2024)
by: Brown, Neil C. C., et al.
Published: (2024)
An Experimental Study of Satisfaction Response: Evaluation of Online Collaborative Learning
by: Cheng, Xusen, et al.
Published: (2023)
by: Cheng, Xusen, et al.
Published: (2023)
Steps Towards an Infrastructure for Scholarly Synthesis
by: Chan, Joel, et al.
Published: (2024)
by: Chan, Joel, et al.
Published: (2024)
Evaluating Authoring Tools with the Explorable Authoring Requirements
by: Salmen, Frederic, et al.
Published: (2024)
by: Salmen, Frederic, et al.
Published: (2024)
Surfacing and Applying Meaning: Supporting Hermeneutical Autonomy for LGBTQ+ People in Taiwan
by: Chen, Yi-Tong, et al.
Published: (2026)
by: Chen, Yi-Tong, et al.
Published: (2026)
Evaluating the Effectiveness of LLMs in Introductory Computer Science Education: A Semester-Long Field Study
by: Lyu, Wenhan, et al.
Published: (2024)
by: Lyu, Wenhan, et al.
Published: (2024)
SportsBuddy: Designing and Evaluating an AI-Powered Sports Video Storytelling Tool Through Real-World Deployment
by: Lin, Tica, et al.
Published: (2025)
by: Lin, Tica, et al.
Published: (2025)
The Role of Task Complexity in Reducing AI Plagiarism: A Study of Generative AI Tools
by: Toker, Sacip, et al.
Published: (2024)
by: Toker, Sacip, et al.
Published: (2024)
Do We Know What They Know We Know? Calibrating Student Trust in AI and Human Responses Through Mutual Theory of Mind
by: Pal, Olivia, et al.
Published: (2026)
by: Pal, Olivia, et al.
Published: (2026)
AI Conversational Tutors in Foreign Language Learning: A Mixed-Methods Evaluation Study
by: Avouris, Nikolaos
Published: (2025)
by: Avouris, Nikolaos
Published: (2025)
Evaluating Generative AI in the Lab: Methodological Challenges and Guidelines
by: Park, Hyerim, et al.
Published: (2026)
by: Park, Hyerim, et al.
Published: (2026)
Open-Source Tool for Evaluating Human-Generated vs. AI-Generated Medical Notes Using the PDQI-9 Framework
by: Sultan, Iyad
Published: (2025)
by: Sultan, Iyad
Published: (2025)
Evaluating Actionability in Explainable AI
by: Mansi, Gennie, et al.
Published: (2026)
by: Mansi, Gennie, et al.
Published: (2026)
Useful for Exploration, Risky for Precision: Evaluating AI Tools in Academic Research
by: Dathe, Anthea, et al.
Published: (2026)
by: Dathe, Anthea, et al.
Published: (2026)
Evaluation of a Provenance Management Tool for Immersive Virtual Fieldwork
by: Bernstetter, Armin, et al.
Published: (2025)
by: Bernstetter, Armin, et al.
Published: (2025)
Reimagining Legal Fact Verification with GenAI: Toward Effective Human-AI Collaboration
by: Han, Sirui, et al.
Published: (2026)
by: Han, Sirui, et al.
Published: (2026)
Network Traffic as a Scalable Ethnographic Lens for Understanding University Students' AI Tool Practices
by: Hu, Donghan, et al.
Published: (2025)
by: Hu, Donghan, et al.
Published: (2025)
Evaluating Effectiveness of Interactivity in Contour-based Geospatial Visualizations
by: Nayeem, Abdullah-Al-Raihan, et al.
Published: (2024)
by: Nayeem, Abdullah-Al-Raihan, et al.
Published: (2024)
AI Academy: Building Generative AI Literacy in Higher Ed Instructors
by: Chen, Si, et al.
Published: (2025)
by: Chen, Si, et al.
Published: (2025)
Studying Self-Care with Generative AI Tools: Lessons for Design
by: Capel, Tara, et al.
Published: (2024)
by: Capel, Tara, et al.
Published: (2024)
Convivial Fabrication: Towards Relational Computational Tools For and From Craft Practices
by: Batra, Ritik, et al.
Published: (2026)
by: Batra, Ritik, et al.
Published: (2026)
Towards Effective Multidisciplinary Health and HCI Teams based on AI Framework
by: Almutairi, Mohammed, et al.
Published: (2025)
by: Almutairi, Mohammed, et al.
Published: (2025)
Designing for Human-Agent Alignment: Understanding what humans want from their agents
by: Goyal, Nitesh, et al.
Published: (2024)
by: Goyal, Nitesh, et al.
Published: (2024)
Rememo: A Research-through-Design Inquiry Towards an AI-in-the-loop Therapist's Tool for Dementia Reminiscence
by: Seah, Celeste, et al.
Published: (2026)
by: Seah, Celeste, et al.
Published: (2026)
Co-Designing Collaborative Generative AI Tools for Freelancers
by: Imteyaz, Kashif, et al.
Published: (2026)
by: Imteyaz, Kashif, et al.
Published: (2026)
Towards Metrics for Evaluating Creativity in Visualisation Design
by: Owen, Aron E, et al.
Published: (2024)
by: Owen, Aron E, et al.
Published: (2024)
Y-AR: A Mixed Reality CAD Tool for 3D Wire Bending
by: Feng, Shuo, et al.
Published: (2024)
by: Feng, Shuo, et al.
Published: (2024)
Toward Human-Centered Human-AI Interaction: Advances in Theoretical Frameworks and Practice
by: Gao, Zaifeng, et al.
Published: (2026)
by: Gao, Zaifeng, et al.
Published: (2026)
From Interaction to Impact: Towards Safer AI Agents Through Understanding and Evaluating Mobile UI Operation Impacts
by: Zhang, Zhuohao Jerry, et al.
Published: (2024)
by: Zhang, Zhuohao Jerry, et al.
Published: (2024)
Similar Items
-
Troubling Taxonomies in GenAI Evaluation
by: Berman, Glen, et al.
Published: (2024) -
The Case for "Thick Evaluations" of Cultural Representation in AI
by: Qadri, Rida, et al.
Published: (2025) -
"It Became My Buddy, But I'm Not Afraid to Disagree": A Multi-Session Study of UX Evaluators Collaborating with Conversational AI Assistants
by: Kuang, Emily, et al.
Published: (2026) -
Farsight: Fostering Responsible AI Awareness During AI Application Prototyping
by: Wang, Zijie J., et al.
Published: (2024) -
AI Trust Reshaping Administrative Burdens: Understanding Trust-Burden Dynamics in LLM-Assisted Benefits Systems
by: Jo, Jeongwon, et al.
Published: (2025)