Healthcare LLM Benchmarks Are Only as Good as Their Explicit Assumptions

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Raman, Naveen, Cortes-Gomez, Santiago, Rubio, Mateo Dulce, Fang, Fei, Wilder, Bryan
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910246390202368
author Raman, Naveen
Cortes-Gomez, Santiago
Rubio, Mateo Dulce
Fang, Fei
Wilder, Bryan
author_facet Raman, Naveen
Cortes-Gomez, Santiago
Rubio, Mateo Dulce
Fang, Fei
Wilder, Bryan
contents Benchmarks are necessary for healthcare evaluation, but are not sufficient for predicting deployment performance. Our position is that the evaluation--deployment gap arises not because of poorly designed benchmarks, but from implicit assumptions about how users interact with models that cannot be surfaced from benchmarks alone. To make this precise, we propose a classification of assumptions into two categories: task, which can be tested from conversation data alone, and outcome, which requires outcome data and behavioral studies for testing. Critically, outcome assumptions depend on human behavior, something that even well-designed benchmarks cannot directly observe. To demonstrate the operationality of this framework, we retrospectively analyze a healthcare RCT as a case study and find that the gap naturally separates into task and outcome gaps of roughly equal size. To address this, we make two contributions: first, we propose BenchmarkCards, an artifact that documents assumptions, and second, we propose staged evaluation, a procedure that systematically tests assumptions and evaluates performance.
format Preprint
id arxiv_https___arxiv_org_abs_2605_22612
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Healthcare LLM Benchmarks Are Only as Good as Their Explicit Assumptions
Raman, Naveen
Cortes-Gomez, Santiago
Rubio, Mateo Dulce
Fang, Fei
Wilder, Bryan
Computers and Society
Artificial Intelligence
Machine Learning
Benchmarks are necessary for healthcare evaluation, but are not sufficient for predicting deployment performance. Our position is that the evaluation--deployment gap arises not because of poorly designed benchmarks, but from implicit assumptions about how users interact with models that cannot be surfaced from benchmarks alone. To make this precise, we propose a classification of assumptions into two categories: task, which can be tested from conversation data alone, and outcome, which requires outcome data and behavioral studies for testing. Critically, outcome assumptions depend on human behavior, something that even well-designed benchmarks cannot directly observe. To demonstrate the operationality of this framework, we retrospectively analyze a healthcare RCT as a case study and find that the gap naturally separates into task and outcome gaps of roughly equal size. To address this, we make two contributions: first, we propose BenchmarkCards, an artifact that documents assumptions, and second, we propose staged evaluation, a procedure that systematically tests assumptions and evaluates performance.
title Healthcare LLM Benchmarks Are Only as Good as Their Explicit Assumptions
topic Computers and Society
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2605.22612