Evaluating LLMs in Medicine: A Call for Rigor, Transparency

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Alwakeel, Mahmoud, Nagori, Aditya, Krishnamoorthy, Vijay, Kamaleswaran, Rishikesan
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915385612173312
author Alwakeel, Mahmoud
Nagori, Aditya
Krishnamoorthy, Vijay
Kamaleswaran, Rishikesan
author_facet Alwakeel, Mahmoud
Nagori, Aditya
Krishnamoorthy, Vijay
Kamaleswaran, Rishikesan
contents Objectives: To evaluate the current limitations of large language models (LLMs) in medical question answering, focusing on the quality of datasets used for their evaluation. Materials and Methods: Widely-used benchmark datasets, including MedQA, MedMCQA, PubMedQA, and MMLU, were reviewed for their rigor, transparency, and relevance to clinical scenarios. Alternatives, such as challenge questions in medical journals, were also analyzed to identify their potential as unbiased evaluation tools. Results: Most existing datasets lack clinical realism, transparency, and robust validation processes. Publicly available challenge questions offer some benefits but are limited by their small size, narrow scope, and exposure to LLM training. These gaps highlight the need for secure, comprehensive, and representative datasets. Conclusion: A standardized framework is critical for evaluating LLMs in medicine. Collaborative efforts among institutions and policymakers are needed to ensure datasets and methodologies are rigorous, unbiased, and reflective of clinical complexities.
format Preprint
id arxiv_https___arxiv_org_abs_2507_08916
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluating LLMs in Medicine: A Call for Rigor, Transparency
Alwakeel, Mahmoud
Nagori, Aditya
Krishnamoorthy, Vijay
Kamaleswaran, Rishikesan
Computation and Language
Objectives: To evaluate the current limitations of large language models (LLMs) in medical question answering, focusing on the quality of datasets used for their evaluation. Materials and Methods: Widely-used benchmark datasets, including MedQA, MedMCQA, PubMedQA, and MMLU, were reviewed for their rigor, transparency, and relevance to clinical scenarios. Alternatives, such as challenge questions in medical journals, were also analyzed to identify their potential as unbiased evaluation tools. Results: Most existing datasets lack clinical realism, transparency, and robust validation processes. Publicly available challenge questions offer some benefits but are limited by their small size, narrow scope, and exposure to LLM training. These gaps highlight the need for secure, comprehensive, and representative datasets. Conclusion: A standardized framework is critical for evaluating LLMs in medicine. Collaborative efforts among institutions and policymakers are needed to ensure datasets and methodologies are rigorous, unbiased, and reflective of clinical complexities.
title Evaluating LLMs in Medicine: A Call for Rigor, Transparency
topic Computation and Language
url https://arxiv.org/abs/2507.08916