Saved in:
Bibliographic Details
Main Authors: Correa-Guillén, Alexis, Gómez-Rodríguez, Carlos, Vilares, David
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2511.15355
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917367706025984
author Correa-Guillén, Alexis
Gómez-Rodríguez, Carlos
Vilares, David
author_facet Correa-Guillén, Alexis
Gómez-Rodríguez, Carlos
Vilares, David
contents We introduce HEAD-QA v2, an expanded and updated version of a Spanish/English healthcare multiple-choice reasoning dataset originally released by Vilares and Gómez-Rodríguez (2019). The update responds to the growing need for high-quality datasets that capture the linguistic and conceptual complexity of healthcare reasoning. We extend the dataset to over 12,000 questions from ten years of Spanish professional exams, benchmark several open-source LLMs using prompting, RAG, and probability-based answer selection, and provide additional multilingual versions to support future work. Results indicate that performance is mainly driven by model scale and intrinsic reasoning ability, with complex inference strategies obtaining limited gains. Together, these results establish HEAD-QA v2 as a reliable resource for advancing research on biomedical reasoning and model improvement.
format Preprint
id arxiv_https___arxiv_org_abs_2511_15355
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HEAD-QA v2: Expanding a Healthcare Benchmark for Reasoning
Correa-Guillén, Alexis
Gómez-Rodríguez, Carlos
Vilares, David
Computation and Language
We introduce HEAD-QA v2, an expanded and updated version of a Spanish/English healthcare multiple-choice reasoning dataset originally released by Vilares and Gómez-Rodríguez (2019). The update responds to the growing need for high-quality datasets that capture the linguistic and conceptual complexity of healthcare reasoning. We extend the dataset to over 12,000 questions from ten years of Spanish professional exams, benchmark several open-source LLMs using prompting, RAG, and probability-based answer selection, and provide additional multilingual versions to support future work. Results indicate that performance is mainly driven by model scale and intrinsic reasoning ability, with complex inference strategies obtaining limited gains. Together, these results establish HEAD-QA v2 as a reliable resource for advancing research on biomedical reasoning and model improvement.
title HEAD-QA v2: Expanding a Healthcare Benchmark for Reasoning
topic Computation and Language
url https://arxiv.org/abs/2511.15355