Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: van der Maden, Willem, Sadek, Malak, Xiao, Ziang, Mottelson, Aske, Liao, Q. Vera, Zhu, Jichen
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910141848223744
author van der Maden, Willem
Sadek, Malak
Xiao, Ziang
Mottelson, Aske
Liao, Q. Vera
Zhu, Jichen
author_facet van der Maden, Willem
Sadek, Malak
Xiao, Ziang
Mottelson, Aske
Liao, Q. Vera
Zhu, Jichen
contents How do product teams evaluate LLM-powered products? As organizations integrate large language models (LLMs) into digital products, their unpredictable nature makes traditional evaluation approaches inadequate, yet little is known about how practitioners navigate this challenge. Through interviews with nineteen practitioners across diverse sectors, we identify ten evaluation practices spanning informal 'vibe checks' to organizational meta-work. Beyond confirming four documented challenges, we introduce a novel fifth we call the results-actionability gap, in which practitioners gather evaluation data but cannot translate findings into concrete improvements. Drawing on patterns from successful teams, we contribute strategies to bridge this gap, supporting practitioners' formalization journey from ad-hoc interpretive practices (e.g., vibe checks) toward systematic evaluation. Our analysis suggests these interpretive practices are necessary adaptations to LLM characteristics rather than methodological failures. For HCI researchers, this presents a research opportunity to support practitioners in systematizing emerging practices rather than developing new evaluation frameworks.
format Preprint
id arxiv_https___arxiv_org_abs_2604_16304
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild
van der Maden, Willem
Sadek, Malak
Xiao, Ziang
Mottelson, Aske
Liao, Q. Vera
Zhu, Jichen
Software Engineering
Artificial Intelligence
Human-Computer Interaction
How do product teams evaluate LLM-powered products? As organizations integrate large language models (LLMs) into digital products, their unpredictable nature makes traditional evaluation approaches inadequate, yet little is known about how practitioners navigate this challenge. Through interviews with nineteen practitioners across diverse sectors, we identify ten evaluation practices spanning informal 'vibe checks' to organizational meta-work. Beyond confirming four documented challenges, we introduce a novel fifth we call the results-actionability gap, in which practitioners gather evaluation data but cannot translate findings into concrete improvements. Drawing on patterns from successful teams, we contribute strategies to bridge this gap, supporting practitioners' formalization journey from ad-hoc interpretive practices (e.g., vibe checks) toward systematic evaluation. Our analysis suggests these interpretive practices are necessary adaptations to LLM characteristics rather than methodological failures. For HCI researchers, this presents a research opportunity to support practitioners in systematizing emerging practices rather than developing new evaluation frameworks.
title Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild
topic Software Engineering
Artificial Intelligence
Human-Computer Interaction
url https://arxiv.org/abs/2604.16304