Designing Staged Evaluation Workflows for LLMs: Integrating Domain Experts, Lay Users, and Model-Generated Evaluation Criteria
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Szymanski, Annalisa, Gebreegziabher, Simret Araya, Anuyah, Oghenemaro, Metoyer, Ronald A., Li, Toby Jia-Jun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Key Considerations for Domain Expert Involvement in LLM Design and Evaluation: An Ethnographic Study
von: Szymanski, Annalisa, et al.
Veröffentlicht: (2026)
von: Szymanski, Annalisa, et al.
Veröffentlicht: (2026)
MultEval: Supporting Collaborative Alignment for LLM-as-a-Judge Evaluation Criteria
von: Chiang, Charles, et al.
Veröffentlicht: (2026)
von: Chiang, Charles, et al.
Veröffentlicht: (2026)
Limitations of the LLM-as-a-Judge Approach for Evaluating LLM Outputs in Expert Knowledge Tasks
von: Szymanski, Annalisa, et al.
Veröffentlicht: (2024)
von: Szymanski, Annalisa, et al.
Veröffentlicht: (2024)
Supporting Co-Adaptive Machine Teaching through Human Concept Learning and Cognitive Theories
von: Gebreegziabher, Simret Araya, et al.
Veröffentlicht: (2024)
von: Gebreegziabher, Simret Araya, et al.
Veröffentlicht: (2024)
Stayin' Aligned Over Time: Towards Longitudinal Human-LLM Alignment via Contextual Reflection and Privacy-Preserving Behavioral Data
von: Gebreegziabher, Simret Araya, et al.
Veröffentlicht: (2026)
von: Gebreegziabher, Simret Araya, et al.
Veröffentlicht: (2026)
Leveraging Variation Theory in Counterfactual Data Augmentation for Optimized Active Learning
von: Gebreegziabher, Simret Araya, et al.
Veröffentlicht: (2024)
von: Gebreegziabher, Simret Araya, et al.
Veröffentlicht: (2024)
A Taxonomy for Human-LLM Interaction Modes: An Initial Exploration
von: Gao, Jie, et al.
Veröffentlicht: (2024)
von: Gao, Jie, et al.
Veröffentlicht: (2024)
The Behavioral Fabric of LLM-Powered GUI Agents: Human Values and Interaction Outcomes
von: Gebreegziabher, Simret Araya, et al.
Veröffentlicht: (2026)
von: Gebreegziabher, Simret Araya, et al.
Veröffentlicht: (2026)
Toward a Human-Centered Evaluation Framework for Trustworthy LLM-Powered GUI Agents
von: Chen, Chaoran, et al.
Veröffentlicht: (2025)
von: Chen, Chaoran, et al.
Veröffentlicht: (2025)
Comparing Human Oversight Strategies for Computer-Use Agents
von: Chen, Chaoran, et al.
Veröffentlicht: (2026)
von: Chen, Chaoran, et al.
Veröffentlicht: (2026)
Online Safety for All: Sociocultural Insights from a Systematic Review of Youth Online Safety in the Global South
von: Oguine, Ozioma C., et al.
Veröffentlicht: (2025)
von: Oguine, Ozioma C., et al.
Veröffentlicht: (2025)
CoCo Matrix: Taxonomy of Cognitive Contributions in Co-writing with Intelligent Agents
von: Wan, Ruyuan, et al.
Veröffentlicht: (2024)
von: Wan, Ruyuan, et al.
Veröffentlicht: (2024)
Bridging the AI Adoption Gap: Designing an Interactive Pedagogical Agent for Higher Education Instructors
von: Chen, Si, et al.
Veröffentlicht: (2025)
von: Chen, Si, et al.
Veröffentlicht: (2025)
Dark Patterns Meet GUI Agents: LLM Agent Susceptibility to Manipulative Interfaces and the Role of Human Oversight
von: Tang, Jingyu, et al.
Veröffentlicht: (2025)
von: Tang, Jingyu, et al.
Veröffentlicht: (2025)
The Obvious Invisible Threat: LLM-Powered GUI Agents' Vulnerability to Fine-Print Injections
von: Chen, Chaoran, et al.
Veröffentlicht: (2025)
von: Chen, Chaoran, et al.
Veröffentlicht: (2025)
ALLOY: Generating Reusable Agent Workflows from User Demonstration
von: Li, Jiawen, et al.
Veröffentlicht: (2025)
von: Li, Jiawen, et al.
Veröffentlicht: (2025)
Flowy: Supporting UX Design Decisions Through AI-Driven Pattern Annotation in Multi-Screen User Flows
von: Lu, Yuwen, et al.
Veröffentlicht: (2024)
von: Lu, Yuwen, et al.
Veröffentlicht: (2024)
Leveraging Large Language Models to Enhance Domain Expert Inclusion in Data Science Workflows
von: Shih, Jasmine Y., et al.
Veröffentlicht: (2024)
von: Shih, Jasmine Y., et al.
Veröffentlicht: (2024)
Designing an LLM-Based Behavioral Activation Chatbot for Young People with Depression: Insights from an Evaluation with Artificial Users and Clinical Experts
von: Kuhlmeier, Florian Onur, et al.
Veröffentlicht: (2025)
von: Kuhlmeier, Florian Onur, et al.
Veröffentlicht: (2025)
Exploring Conversational Design Choices in LLMs for Pedagogical Purposes: Socratic and Narrative Approaches for Improving Instructor's Teaching Practice
von: Chen, Si, et al.
Veröffentlicht: (2025)
von: Chen, Si, et al.
Veröffentlicht: (2025)
From Awareness to Action: Exploring End-User Empowerment Interventions for Dark Patterns in UX
von: Lu, Yuwen, et al.
Veröffentlicht: (2023)
von: Lu, Yuwen, et al.
Veröffentlicht: (2023)
Algorithm, Expert, or Both? Evaluating the Role of Feature Selection Methods on User Preferences and Reliance
von: Kornowicz, Jaroslaw, et al.
Veröffentlicht: (2024)
von: Kornowicz, Jaroslaw, et al.
Veröffentlicht: (2024)
CollabCoder: A Lower-barrier, Rigorous Workflow for Inductive Collaborative Qualitative Analysis with Large Language Models
von: Gao, Jie, et al.
Veröffentlicht: (2023)
von: Gao, Jie, et al.
Veröffentlicht: (2023)
LLMs are the Ideal Candidate for Mixed-Initiative Game Design Pillar Workflows
von: Geheeb, Julian, et al.
Veröffentlicht: (2026)
von: Geheeb, Julian, et al.
Veröffentlicht: (2026)
How can LLMs Support Policy Researchers? Evaluating an LLM-Assisted Workflow for Large-Scale Unstructured Data
von: Liu, Yuhan, et al.
Veröffentlicht: (2026)
von: Liu, Yuhan, et al.
Veröffentlicht: (2026)
AI Academy: Building Generative AI Literacy in Higher Ed Instructors
von: Chen, Si, et al.
Veröffentlicht: (2025)
von: Chen, Si, et al.
Veröffentlicht: (2025)
Sketchar: Supporting Character Design and Illustration Prototyping Using Generative AI
von: Ling, Long, et al.
Veröffentlicht: (2025)
von: Ling, Long, et al.
Veröffentlicht: (2025)
If You Had to Pitch Your Ideal Software -- Evaluating Large Language Models to Support User Scenario Writing for User Experience Experts and Laypersons
von: Stadler, Patrick, et al.
Veröffentlicht: (2025)
von: Stadler, Patrick, et al.
Veröffentlicht: (2025)
Designing and Evaluating Scalable Privacy Awareness and Control User Interfaces for Mixed Reality
von: Strauss, Marvin, et al.
Veröffentlicht: (2024)
von: Strauss, Marvin, et al.
Veröffentlicht: (2024)
Towards a Design Guideline for RPA Evaluation: A Survey of Large Language Model-Based Role-Playing Agents
von: Chen, Chaoran, et al.
Veröffentlicht: (2025)
von: Chen, Chaoran, et al.
Veröffentlicht: (2025)
Building AI Literacy at Home: How Families Navigate Children's Self-Directed Learning with AI
von: Xie, Jingyi, et al.
Veröffentlicht: (2025)
von: Xie, Jingyi, et al.
Veröffentlicht: (2025)
A Systematic Review of User-Centred Evaluation of Explainable AI in Healthcare
von: Donoso-Guzmán, Ivania, et al.
Veröffentlicht: (2025)
von: Donoso-Guzmán, Ivania, et al.
Veröffentlicht: (2025)
Experience Level Influences User's Criteria for Avatar Animation Realism
von: Huang, Yudong, et al.
Veröffentlicht: (2025)
von: Huang, Yudong, et al.
Veröffentlicht: (2025)
EvalLM: Interactive Evaluation of Large Language Model Prompts on User-Defined Criteria
von: Kim, Tae Soo, et al.
Veröffentlicht: (2023)
von: Kim, Tae Soo, et al.
Veröffentlicht: (2023)
When LLMs fall short in Deductive Coding: Model Comparison and Human AI Collaboration Workflow Design
von: Li, Zijian, et al.
Veröffentlicht: (2025)
von: Li, Zijian, et al.
Veröffentlicht: (2025)
Guiding, not Driving: Design and Evaluation of a Command-Based User Interface for Teleoperation of Autonomous Vehicles
von: Tener, Felix, et al.
Veröffentlicht: (2025)
von: Tener, Felix, et al.
Veröffentlicht: (2025)
Implications of AI Involvement for Trust in Expert Advisory Workflows Under Epistemic Dependence
von: Kim, Dennis, et al.
Veröffentlicht: (2026)
von: Kim, Dennis, et al.
Veröffentlicht: (2026)
Co-persona: Leveraging LLMs and Expert Collaboration to Understand User Personas through Social Media Data Analysis
von: Yin, Min, et al.
Veröffentlicht: (2025)
von: Yin, Min, et al.
Veröffentlicht: (2025)
Exploring Trust Calibration in XAI - The Impact of Exposing Model Limitations to Lay Users
von: Ventura, Alfio, et al.
Veröffentlicht: (2026)
von: Ventura, Alfio, et al.
Veröffentlicht: (2026)
Visualizing Historical Book Trade Data: An Iterative Design Study with Close Collaboration with Domain Experts
von: Xing, Yiwen, et al.
Veröffentlicht: (2023)
von: Xing, Yiwen, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Key Considerations for Domain Expert Involvement in LLM Design and Evaluation: An Ethnographic Study
von: Szymanski, Annalisa, et al.
Veröffentlicht: (2026) -
MultEval: Supporting Collaborative Alignment for LLM-as-a-Judge Evaluation Criteria
von: Chiang, Charles, et al.
Veröffentlicht: (2026) -
Limitations of the LLM-as-a-Judge Approach for Evaluating LLM Outputs in Expert Knowledge Tasks
von: Szymanski, Annalisa, et al.
Veröffentlicht: (2024) -
Supporting Co-Adaptive Machine Teaching through Human Concept Learning and Cognitive Theories
von: Gebreegziabher, Simret Araya, et al.
Veröffentlicht: (2024) -
Stayin' Aligned Over Time: Towards Longitudinal Human-LLM Alignment via Contextual Reflection and Privacy-Preserving Behavioral Data
von: Gebreegziabher, Simret Araya, et al.
Veröffentlicht: (2026)