Overview of the MedHopQA track at BioCreative IX: track description, participation and evaluation of systems for multi-hop medical question answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Islamaj, Rezarta, Chan, Joey, Leaman, Robert, Jung, Jongmyung, Hwang, Hyeongsoon, Nguyen, Quoc-An, Le, Hoang-Quynh, Saisudha, Harikrishnan Gurushankar, Chandrasekar, Ganesh, Taktashov, Rustam R., Bizyukova, Nadezhda Yu., Conceição, Sofia I. R., Lopes, Paulo R. C., Salam, Reem Abdel, Adewunmi, Mary, Lu, Zhiyong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914559723307008
author Islamaj, Rezarta
Chan, Joey
Leaman, Robert
Jung, Jongmyung
Hwang, Hyeongsoon
Nguyen, Quoc-An
Le, Hoang-Quynh
Saisudha, Harikrishnan Gurushankar
Chandrasekar, Ganesh
Taktashov, Rustam R.
Bizyukova, Nadezhda Yu.
Conceição, Sofia I. R.
Lopes, Paulo R. C.
Salam, Reem Abdel
Adewunmi, Mary
Lu, Zhiyong
author_facet Islamaj, Rezarta
Chan, Joey
Leaman, Robert
Jung, Jongmyung
Hwang, Hyeongsoon
Nguyen, Quoc-An
Le, Hoang-Quynh
Saisudha, Harikrishnan Gurushankar
Chandrasekar, Ganesh
Taktashov, Rustam R.
Bizyukova, Nadezhda Yu.
Conceição, Sofia I. R.
Lopes, Paulo R. C.
Salam, Reem Abdel
Adewunmi, Mary
Lu, Zhiyong
contents Multi-hop question answering (QA) remains a significant challenge in the biomedical domain, requiring systems to integrate information across multiple sources to answer complex questions. To address this problem, the BioCreative IX MedHopQA shared task was designed to benchmark in multi-hop reasoning for large language models (LLMs). We developed a novel dataset of 1,000 challenging QA pairs spanning diseases, genes, and chemicals, with particular emphasis on rare diseases. Each question was constructed to require two-hop reasoning through the integration of information from two distinct Wikipedia pages. The challenge attracted 48 submissions from 13 teams. Systems were evaluated using both surface string comparison and conceptual accuracy (MedCPT score). The results showed a substantial performance gap between baseline LLMs and enhanced systems. The top-ranked submission achieved an 89.30% F1 score on the MedCPT metric and an 87.30% exact match (EM) score, compared with 67.40% and 60.20%, respectively, for the zero-shot baseline. A central finding of the challenge was that retrieval-augmented generation (RAG) and related retrieval-based strategies were critical for strong performance. In addition, concept-level evaluation improved answer assessment when correct responses differed in surface form. The MedHopQA dataset is publicly available to support continued progress in this important area. Challenge materials: https://www.ncbi.nlm.nih.gov/research/bionlp/medhopqa and benchmark https://www.codabench.org/competitions/7609/
format Preprint
id arxiv_https___arxiv_org_abs_2605_12313
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Overview of the MedHopQA track at BioCreative IX: track description, participation and evaluation of systems for multi-hop medical question answering
Islamaj, Rezarta
Chan, Joey
Leaman, Robert
Jung, Jongmyung
Hwang, Hyeongsoon
Nguyen, Quoc-An
Le, Hoang-Quynh
Saisudha, Harikrishnan Gurushankar
Chandrasekar, Ganesh
Taktashov, Rustam R.
Bizyukova, Nadezhda Yu.
Conceição, Sofia I. R.
Lopes, Paulo R. C.
Salam, Reem Abdel
Adewunmi, Mary
Lu, Zhiyong
Computation and Language
Information Retrieval
Multi-hop question answering (QA) remains a significant challenge in the biomedical domain, requiring systems to integrate information across multiple sources to answer complex questions. To address this problem, the BioCreative IX MedHopQA shared task was designed to benchmark in multi-hop reasoning for large language models (LLMs). We developed a novel dataset of 1,000 challenging QA pairs spanning diseases, genes, and chemicals, with particular emphasis on rare diseases. Each question was constructed to require two-hop reasoning through the integration of information from two distinct Wikipedia pages. The challenge attracted 48 submissions from 13 teams. Systems were evaluated using both surface string comparison and conceptual accuracy (MedCPT score). The results showed a substantial performance gap between baseline LLMs and enhanced systems. The top-ranked submission achieved an 89.30% F1 score on the MedCPT metric and an 87.30% exact match (EM) score, compared with 67.40% and 60.20%, respectively, for the zero-shot baseline. A central finding of the challenge was that retrieval-augmented generation (RAG) and related retrieval-based strategies were critical for strong performance. In addition, concept-level evaluation improved answer assessment when correct responses differed in surface form. The MedHopQA dataset is publicly available to support continued progress in this important area. Challenge materials: https://www.ncbi.nlm.nih.gov/research/bionlp/medhopqa and benchmark https://www.codabench.org/competitions/7609/
title Overview of the MedHopQA track at BioCreative IX: track description, participation and evaluation of systems for multi-hop medical question answering
topic Computation and Language
Information Retrieval
url https://arxiv.org/abs/2605.12313