Key Considerations for Domain Expert Involvement in LLM Design and Evaluation: An Ethnographic Study
Fuente:
arXiv
Salvato in:
| Autori principali: | Szymanski, Annalisa, Anuyah, Oghenemaro, Li, Toby Jia-Jun, Metoyer, Ronald A. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Designing Staged Evaluation Workflows for LLMs: Integrating Domain Experts, Lay Users, and Model-Generated Evaluation Criteria
di: Szymanski, Annalisa, et al.
Pubblicazione: (2024)
di: Szymanski, Annalisa, et al.
Pubblicazione: (2024)
Limitations of the LLM-as-a-Judge Approach for Evaluating LLM Outputs in Expert Knowledge Tasks
di: Szymanski, Annalisa, et al.
Pubblicazione: (2024)
di: Szymanski, Annalisa, et al.
Pubblicazione: (2024)
Designing and Evaluating Malinowski's Lens: An AI-Native Educational Game for Ethnographic Learning
di: Hoffmann, Michael, et al.
Pubblicazione: (2025)
di: Hoffmann, Michael, et al.
Pubblicazione: (2025)
From Verification Burden to Trusted Collaboration: Design Goals for LLM-Assisted Literature Reviews
di: Nogueira, Brenda, et al.
Pubblicazione: (2025)
di: Nogueira, Brenda, et al.
Pubblicazione: (2025)
Exploring Conversational Design Choices in LLMs for Pedagogical Purposes: Socratic and Narrative Approaches for Improving Instructor's Teaching Practice
di: Chen, Si, et al.
Pubblicazione: (2025)
di: Chen, Si, et al.
Pubblicazione: (2025)
Luminate: Structured Generation and Exploration of Design Space with Large Language Models for Human-AI Co-Creation
di: Suh, Sangho, et al.
Pubblicazione: (2023)
di: Suh, Sangho, et al.
Pubblicazione: (2023)
LLM Enhancement with Domain Expert Mental Model to Reduce LLM Hallucination with Causal Prompt Engineering
di: Kovalerchuk, Boris, et al.
Pubblicazione: (2025)
di: Kovalerchuk, Boris, et al.
Pubblicazione: (2025)
Human-Centered Evaluation of an LLM-Based Process Modeling Copilot: A Mixed-Methods Study with Domain Experts
di: Lauer, Chantale, et al.
Pubblicazione: (2026)
di: Lauer, Chantale, et al.
Pubblicazione: (2026)
PersonaFlow: Designing LLM-Simulated Expert Perspectives for Enhanced Research Ideation
di: Liu, Yiren, et al.
Pubblicazione: (2024)
di: Liu, Yiren, et al.
Pubblicazione: (2024)
Beyond Permissions: Investigating Mobile Personalization with Simulated Personas
di: Khalilov, Ibrahim, et al.
Pubblicazione: (2025)
di: Khalilov, Ibrahim, et al.
Pubblicazione: (2025)
Synthetic Interlocutors. Experiments with Generative AI to Prolong Ethnographic Encounters
di: Søltoft, Johan Irving, et al.
Pubblicazione: (2024)
di: Søltoft, Johan Irving, et al.
Pubblicazione: (2024)
Creativity as a Human Right: Design Considerations for Computational Creativity Systems
di: Issak, Alayt
Pubblicazione: (2025)
di: Issak, Alayt
Pubblicazione: (2025)
LLARS: Enabling Domain Expert & Developer Collaboration for LLM Prompting, Generation and Evaluation
di: Steigerwald, Philipp, et al.
Pubblicazione: (2026)
di: Steigerwald, Philipp, et al.
Pubblicazione: (2026)
Tempo: Helping Data Scientists and Domain Experts Collaboratively Specify Predictive Modeling Tasks
di: Sivaraman, Venkatesh, et al.
Pubblicazione: (2025)
di: Sivaraman, Venkatesh, et al.
Pubblicazione: (2025)
Dark Patterns Meet GUI Agents: LLM Agent Susceptibility to Manipulative Interfaces and the Role of Human Oversight
di: Tang, Jingyu, et al.
Pubblicazione: (2025)
di: Tang, Jingyu, et al.
Pubblicazione: (2025)
An Explanatory Model Steering System for Collaboration between Domain Experts and AI
di: Bhattacharya, Aditya, et al.
Pubblicazione: (2024)
di: Bhattacharya, Aditya, et al.
Pubblicazione: (2024)
Towards Real-World Validity in Generative AI Benchmarks: Understanding and Designing Domain-Centered Evaluations for Journalism Practitioners
di: Li, Charlotte, et al.
Pubblicazione: (2025)
di: Li, Charlotte, et al.
Pubblicazione: (2025)
CoCo Matrix: Taxonomy of Cognitive Contributions in Co-writing with Intelligent Agents
di: Wan, Ruyuan, et al.
Pubblicazione: (2024)
di: Wan, Ruyuan, et al.
Pubblicazione: (2024)
Evaluating Human Trust in LLM-Based Planners: A Preliminary Study
di: Chen, Shenghui, et al.
Pubblicazione: (2025)
di: Chen, Shenghui, et al.
Pubblicazione: (2025)
Communication Styles and Reader Preferences of LLM and Human Experts in Explaining Health Information
di: Zhou, Jiawei, et al.
Pubblicazione: (2025)
di: Zhou, Jiawei, et al.
Pubblicazione: (2025)
Data Therapist: Eliciting Domain Knowledge from Subject Matter Experts Using Large Language Models
di: Shin, Sungbok, et al.
Pubblicazione: (2025)
di: Shin, Sungbok, et al.
Pubblicazione: (2025)
Explanatory Debiasing: Involving Domain Experts in the Data Generation Process to Mitigate Representation Bias in AI Systems
di: Bhattacharya, Aditya, et al.
Pubblicazione: (2024)
di: Bhattacharya, Aditya, et al.
Pubblicazione: (2024)
Qualitative Evaluation of LLM-Designed GUI
di: Sawicki, Bartosz, et al.
Pubblicazione: (2026)
di: Sawicki, Bartosz, et al.
Pubblicazione: (2026)
ChartDesign: Towards LLM Designer of Data Visualization
di: Ansari, Mohammed Afaan, et al.
Pubblicazione: (2026)
di: Ansari, Mohammed Afaan, et al.
Pubblicazione: (2026)
Exploring Collaborative GenAI Agents in Synchronous Group Settings: Eliciting Team Perceptions and Design Considerations for the Future of Work
di: Johnson, Janet G., et al.
Pubblicazione: (2025)
di: Johnson, Janet G., et al.
Pubblicazione: (2025)
Vital Insight: Assisting Experts' Context-Driven Sensemaking of Multi-modal Personal Tracking Data Using Visualization and Human-In-The-Loop LLM
di: Li, Jiachen, et al.
Pubblicazione: (2024)
di: Li, Jiachen, et al.
Pubblicazione: (2024)
Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing
di: Jafari, Kiana, et al.
Pubblicazione: (2026)
di: Jafari, Kiana, et al.
Pubblicazione: (2026)
Towards Human-AI Deliberation: Design and Evaluation of LLM-Empowered Deliberative AI for AI-Assisted Decision-Making
di: Ma, Shuai, et al.
Pubblicazione: (2024)
di: Ma, Shuai, et al.
Pubblicazione: (2024)
MultEval: Supporting Collaborative Alignment for LLM-as-a-Judge Evaluation Criteria
di: Chiang, Charles, et al.
Pubblicazione: (2026)
di: Chiang, Charles, et al.
Pubblicazione: (2026)
Dynamik: Syntactically-Driven Dynamic Font Sizing for Emphasis of Key Information
di: Nishida, Naoto, et al.
Pubblicazione: (2025)
di: Nishida, Naoto, et al.
Pubblicazione: (2025)
UICrit: Enhancing Automated Design Evaluation with a UICritique Dataset
di: Duan, Peitong, et al.
Pubblicazione: (2024)
di: Duan, Peitong, et al.
Pubblicazione: (2024)
Predictive Prototyping: Evaluating Design Concepts with ChatGPT
di: Yong, Hilsann, et al.
Pubblicazione: (2026)
di: Yong, Hilsann, et al.
Pubblicazione: (2026)
ALLOY: Generating Reusable Agent Workflows from User Demonstration
di: Li, Jiawen, et al.
Pubblicazione: (2025)
di: Li, Jiawen, et al.
Pubblicazione: (2025)
BlenderLLM: Training Large Language Models for Computer-Aided Design with Self-improvement
di: Du, Yuhao, et al.
Pubblicazione: (2024)
di: Du, Yuhao, et al.
Pubblicazione: (2024)
Mitigating LLM Hallucinations with Knowledge Graphs: A Case Study
di: Li, Harry, et al.
Pubblicazione: (2025)
di: Li, Harry, et al.
Pubblicazione: (2025)
Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences
di: Shankar, Shreya, et al.
Pubblicazione: (2024)
di: Shankar, Shreya, et al.
Pubblicazione: (2024)
LLM-Assisted Knowledge Graph Completion for Curriculum and Domain Modelling in Personalized Higher Education Recommendations
di: Abu-Rasheed, Hasan, et al.
Pubblicazione: (2025)
di: Abu-Rasheed, Hasan, et al.
Pubblicazione: (2025)
AI Expert Twin: Capturing Expert Cognition for Human-Centred, Practice-Based Learning
di: Yuan, Annie, et al.
Pubblicazione: (2026)
di: Yuan, Annie, et al.
Pubblicazione: (2026)
Towards Considerate Human-Robot Coexistence: A Dual-Space Framework of Robot Design and Human Perception in Healthcare
di: Bai, Yuanchen, et al.
Pubblicazione: (2026)
di: Bai, Yuanchen, et al.
Pubblicazione: (2026)
LLM Bazaar: A Service Design for Supporting Collaborative Learning with an LLM-Powered Multi-Party Collaboration Infrastructure
di: Wu, Zhen, et al.
Pubblicazione: (2025)
di: Wu, Zhen, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Designing Staged Evaluation Workflows for LLMs: Integrating Domain Experts, Lay Users, and Model-Generated Evaluation Criteria
di: Szymanski, Annalisa, et al.
Pubblicazione: (2024) -
Limitations of the LLM-as-a-Judge Approach for Evaluating LLM Outputs in Expert Knowledge Tasks
di: Szymanski, Annalisa, et al.
Pubblicazione: (2024) -
Designing and Evaluating Malinowski's Lens: An AI-Native Educational Game for Ethnographic Learning
di: Hoffmann, Michael, et al.
Pubblicazione: (2025) -
From Verification Burden to Trusted Collaboration: Design Goals for LLM-Assisted Literature Reviews
di: Nogueira, Brenda, et al.
Pubblicazione: (2025) -
Exploring Conversational Design Choices in LLMs for Pedagogical Purposes: Socratic and Narrative Approaches for Improving Instructor's Teaching Practice
di: Chen, Si, et al.
Pubblicazione: (2025)