Medical Large Language Model Benchmarks Should Prioritize Construct Validity

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Alaa, Ahmed, Hartvigsen, Thomas, Golchini, Niloufar, Dutta, Shiladitya, Dean, Frances, Raji, Inioluwa Deborah, Zack, Travis
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929759296946176
author Alaa, Ahmed
Hartvigsen, Thomas
Golchini, Niloufar
Dutta, Shiladitya
Dean, Frances
Raji, Inioluwa Deborah
Zack, Travis
author_facet Alaa, Ahmed
Hartvigsen, Thomas
Golchini, Niloufar
Dutta, Shiladitya
Dean, Frances
Raji, Inioluwa Deborah
Zack, Travis
contents Medical large language models (LLMs) research often makes bold claims, from encoding clinical knowledge to reasoning like a physician. These claims are usually backed by evaluation on competitive benchmarks; a tradition inherited from mainstream machine learning. But how do we separate real progress from a leaderboard flex? Medical LLM benchmarks, much like those in other fields, are arbitrarily constructed using medical licensing exam questions. For these benchmarks to truly measure progress, they must accurately capture the real-world tasks they aim to represent. In this position paper, we argue that medical LLM benchmarks should (and indeed can) be empirically evaluated for their construct validity. In the psychological testing literature, "construct validity" refers to the ability of a test to measure an underlying "construct", that is the actual conceptual target of evaluation. By drawing an analogy between LLM benchmarks and psychological tests, we explain how frameworks from this field can provide empirical foundations for validating benchmarks. To put these ideas into practice, we use real-world clinical data in proof-of-concept experiments to evaluate popular medical LLM benchmarks and report significant gaps in their construct validity. Finally, we outline a vision for a new ecosystem of medical LLM evaluation centered around the creation of valid benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2503_10694
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Medical Large Language Model Benchmarks Should Prioritize Construct Validity
Alaa, Ahmed
Hartvigsen, Thomas
Golchini, Niloufar
Dutta, Shiladitya
Dean, Frances
Raji, Inioluwa Deborah
Zack, Travis
Computation and Language
Medical large language models (LLMs) research often makes bold claims, from encoding clinical knowledge to reasoning like a physician. These claims are usually backed by evaluation on competitive benchmarks; a tradition inherited from mainstream machine learning. But how do we separate real progress from a leaderboard flex? Medical LLM benchmarks, much like those in other fields, are arbitrarily constructed using medical licensing exam questions. For these benchmarks to truly measure progress, they must accurately capture the real-world tasks they aim to represent. In this position paper, we argue that medical LLM benchmarks should (and indeed can) be empirically evaluated for their construct validity. In the psychological testing literature, "construct validity" refers to the ability of a test to measure an underlying "construct", that is the actual conceptual target of evaluation. By drawing an analogy between LLM benchmarks and psychological tests, we explain how frameworks from this field can provide empirical foundations for validating benchmarks. To put these ideas into practice, we use real-world clinical data in proof-of-concept experiments to evaluate popular medical LLM benchmarks and report significant gaps in their construct validity. Finally, we outline a vision for a new ecosystem of medical LLM evaluation centered around the creation of valid benchmarks.
title Medical Large Language Model Benchmarks Should Prioritize Construct Validity
topic Computation and Language
url https://arxiv.org/abs/2503.10694