DeVisE: Behavioral Testing of Medical Large Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Tagliabue, Camila Zurdo, Boll, Heloisa Oss, Erdem, Aykut, Erdem, Erkut, Calixto, Iacer
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908852700577792
author Tagliabue, Camila Zurdo
Boll, Heloisa Oss
Erdem, Aykut
Erdem, Erkut
Calixto, Iacer
author_facet Tagliabue, Camila Zurdo
Boll, Heloisa Oss
Erdem, Aykut
Erdem, Erkut
Calixto, Iacer
contents Large language models (LLMs) are increasingly applied in clinical decision support, yet current evaluations rarely reveal whether their outputs reflect genuine medical reasoning or superficial correlations. We introduce DeVisE (Demographics and Vital signs Evaluation), a behavioral testing framework that probes fine-grained clinical understanding through controlled counterfactuals. Using intensive care unit (ICU) discharge notes from MIMIC-IV, we construct both raw (real-world) and template-based (synthetic) variants with single-variable perturbations in demographic (age, gender, ethnicity) and vital sign attributes. We evaluate eight LLMs, spanning general-purpose and medical variants, under zero-shot setting. Model behavior is analyzed through (1) input-level sensitivity, capturing how counterfactuals alter perplexity, and (2) downstream reasoning, measuring their effect on predicted ICU length-of-stay and mortality. Overall, our results show that standard task metrics obscure clinically relevant differences in model behavior, with models differing substantially in how consistently and proportionally they adjust predictions to counterfactual perturbations.
format Preprint
id arxiv_https___arxiv_org_abs_2506_15339
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DeVisE: Behavioral Testing of Medical Large Language Models
Tagliabue, Camila Zurdo
Boll, Heloisa Oss
Erdem, Aykut
Erdem, Erkut
Calixto, Iacer
Computation and Language
Large language models (LLMs) are increasingly applied in clinical decision support, yet current evaluations rarely reveal whether their outputs reflect genuine medical reasoning or superficial correlations. We introduce DeVisE (Demographics and Vital signs Evaluation), a behavioral testing framework that probes fine-grained clinical understanding through controlled counterfactuals. Using intensive care unit (ICU) discharge notes from MIMIC-IV, we construct both raw (real-world) and template-based (synthetic) variants with single-variable perturbations in demographic (age, gender, ethnicity) and vital sign attributes. We evaluate eight LLMs, spanning general-purpose and medical variants, under zero-shot setting. Model behavior is analyzed through (1) input-level sensitivity, capturing how counterfactuals alter perplexity, and (2) downstream reasoning, measuring their effect on predicted ICU length-of-stay and mortality. Overall, our results show that standard task metrics obscure clinically relevant differences in model behavior, with models differing substantially in how consistently and proportionally they adjust predictions to counterfactual perturbations.
title DeVisE: Behavioral Testing of Medical Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2506.15339