MedHopQA: A Disease-Centered Multi-Hop Reasoning Benchmark and Evaluation Framework for LLM-Based Biomedical Question Answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Islamaj, Rezarta, Leaman, Robert, Chan, Joey, Wan, Nicholas, Jin, Qiao, Xie, Natalie, Wilbur, John, Tian, Shubo, Yeganova, Lana, Lai, Po-Ting, Wei, Chih-Hsuan, Yang, Yifan, Ge, Yao, Zhu, Qingqing, Wang, Zhizheng, Lu, Zhiyong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910213132517376
author Islamaj, Rezarta
Leaman, Robert
Chan, Joey
Wan, Nicholas
Jin, Qiao
Xie, Natalie
Wilbur, John
Tian, Shubo
Yeganova, Lana
Lai, Po-Ting
Wei, Chih-Hsuan
Yang, Yifan
Ge, Yao
Zhu, Qingqing
Wang, Zhizheng
Lu, Zhiyong
author_facet Islamaj, Rezarta
Leaman, Robert
Chan, Joey
Wan, Nicholas
Jin, Qiao
Xie, Natalie
Wilbur, John
Tian, Shubo
Yeganova, Lana
Lai, Po-Ting
Wei, Chih-Hsuan
Yang, Yifan
Ge, Yao
Zhu, Qingqing
Wang, Zhizheng
Lu, Zhiyong
contents Evaluating large language models (LLMs) in the biomedical domain requires benchmarks that can distinguish reasoning from pattern matching and remain discriminative as model capabilities improve. Existing biomedical question answering (QA) benchmarks are limited in this respect. Multiple-choice formats can allow models to succeed through answer elimination rather than inference, while widely circulated exam-style datasets are increasingly vulnerable to performance saturation and training data contamination. Multi-hop reasoning, defined as the ability to integrate information across multiple sources to derive an answer, is central to clinically meaningful tasks such as diagnostic support, literature-based discovery, and hypothesis generation, yet remains underrepresented in current biomedical QA benchmarks. MedHopQA is a disease-centered multi-hop reasoning benchmark consisting of 1,000 expert-curated question-answer pairs introduced as a shared task at BioCreative IX. Each question requires synthesis of information across two distinct Wikipedia articles, and answers are provided in an open-ended free-text format. Gold annotations are augmented with ontology-grounded synonym sets from MONDO, NCBI Gene, and NCBI Taxonomy to support both lexical and concept-level evaluation. MedHopQA was constructed through a structured process combining human annotation, triage, iterative verification, and LLM-as-a-judge validation. To reduce leaderboard gaming and contamination risk, the 1,000 scored questions are embedded within a publicly downloadable set of 10,000 questions, with answers withheld, on a CodaBench leaderboard. MedHopQA provides both a benchmark and a reusable framework for constructing future biomedical QA datasets that prioritize compositional reasoning, saturation resistance, and contamination resistance as core design constraints.
format Preprint
id arxiv_https___arxiv_org_abs_2605_12361
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MedHopQA: A Disease-Centered Multi-Hop Reasoning Benchmark and Evaluation Framework for LLM-Based Biomedical Question Answering
Islamaj, Rezarta
Leaman, Robert
Chan, Joey
Wan, Nicholas
Jin, Qiao
Xie, Natalie
Wilbur, John
Tian, Shubo
Yeganova, Lana
Lai, Po-Ting
Wei, Chih-Hsuan
Yang, Yifan
Ge, Yao
Zhu, Qingqing
Wang, Zhizheng
Lu, Zhiyong
Computation and Language
Artificial Intelligence
Information Retrieval
Evaluating large language models (LLMs) in the biomedical domain requires benchmarks that can distinguish reasoning from pattern matching and remain discriminative as model capabilities improve. Existing biomedical question answering (QA) benchmarks are limited in this respect. Multiple-choice formats can allow models to succeed through answer elimination rather than inference, while widely circulated exam-style datasets are increasingly vulnerable to performance saturation and training data contamination. Multi-hop reasoning, defined as the ability to integrate information across multiple sources to derive an answer, is central to clinically meaningful tasks such as diagnostic support, literature-based discovery, and hypothesis generation, yet remains underrepresented in current biomedical QA benchmarks. MedHopQA is a disease-centered multi-hop reasoning benchmark consisting of 1,000 expert-curated question-answer pairs introduced as a shared task at BioCreative IX. Each question requires synthesis of information across two distinct Wikipedia articles, and answers are provided in an open-ended free-text format. Gold annotations are augmented with ontology-grounded synonym sets from MONDO, NCBI Gene, and NCBI Taxonomy to support both lexical and concept-level evaluation. MedHopQA was constructed through a structured process combining human annotation, triage, iterative verification, and LLM-as-a-judge validation. To reduce leaderboard gaming and contamination risk, the 1,000 scored questions are embedded within a publicly downloadable set of 10,000 questions, with answers withheld, on a CodaBench leaderboard. MedHopQA provides both a benchmark and a reusable framework for constructing future biomedical QA datasets that prioritize compositional reasoning, saturation resistance, and contamination resistance as core design constraints.
title MedHopQA: A Disease-Centered Multi-Hop Reasoning Benchmark and Evaluation Framework for LLM-Based Biomedical Question Answering
topic Computation and Language
Artificial Intelligence
Information Retrieval
url https://arxiv.org/abs/2605.12361