EpiQAL: Benchmarking Large Language Models in Epidemiological Question Answering and Reasoning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wei, Mingyang, Min, Dehai, Liu, Zewen, Xie, Yuzhang, Wu, Guanchen, Zhang, Ziyang, Yang, Carl, Lau, Max S. Y., He, Qi, Cheng, Lu, Jin, Wei
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918524416425984
author Wei, Mingyang
Min, Dehai
Liu, Zewen
Xie, Yuzhang
Wu, Guanchen
Zhang, Ziyang
Yang, Carl
Lau, Max S. Y.
He, Qi
Cheng, Lu
Jin, Wei
author_facet Wei, Mingyang
Min, Dehai
Liu, Zewen
Xie, Yuzhang
Wu, Guanchen
Zhang, Ziyang
Yang, Carl
Lau, Max S. Y.
He, Qi
Cheng, Lu
Jin, Wei
contents Reliable epidemiological reasoning requires synthesizing study evidence to infer disease burden, transmission dynamics, and intervention effects at the population level. Existing medical question answering benchmarks primarily emphasize clinical knowledge or patient-level reasoning, yet few systematically evaluate evidence-grounded epidemiological inference. We present EpiQAL, the first diagnostic benchmark for epidemiological question answering across diverse diseases, comprising three subsets built from open-access literature. The three subsets progressively test factual recall, multi-step inference, and conclusion reconstruction under incomplete information, and are constructed through a quality-controlled pipeline combining taxonomy guidance, multi-model verification, and difficulty screening. Experiments on fifteen models spanning open-source and proprietary systems reveal that current LLMs show limited performance on epidemiological reasoning, with multi-step inference posing the greatest challenge. Model rankings shift across subsets, and scale alone does not predict success. Chain-of-Thought prompting benefits multi-step inference but yields mixed results elsewhere. EpiQAL provides fine-grained diagnostic signals for evidence-grounding, inferential reasoning, and conclusion reconstruction.
format Preprint
id arxiv_https___arxiv_org_abs_2601_03471
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle EpiQAL: Benchmarking Large Language Models in Epidemiological Question Answering and Reasoning
Wei, Mingyang
Min, Dehai
Liu, Zewen
Xie, Yuzhang
Wu, Guanchen
Zhang, Ziyang
Yang, Carl
Lau, Max S. Y.
He, Qi
Cheng, Lu
Jin, Wei
Computation and Language
Artificial Intelligence
Reliable epidemiological reasoning requires synthesizing study evidence to infer disease burden, transmission dynamics, and intervention effects at the population level. Existing medical question answering benchmarks primarily emphasize clinical knowledge or patient-level reasoning, yet few systematically evaluate evidence-grounded epidemiological inference. We present EpiQAL, the first diagnostic benchmark for epidemiological question answering across diverse diseases, comprising three subsets built from open-access literature. The three subsets progressively test factual recall, multi-step inference, and conclusion reconstruction under incomplete information, and are constructed through a quality-controlled pipeline combining taxonomy guidance, multi-model verification, and difficulty screening. Experiments on fifteen models spanning open-source and proprietary systems reveal that current LLMs show limited performance on epidemiological reasoning, with multi-step inference posing the greatest challenge. Model rankings shift across subsets, and scale alone does not predict success. Chain-of-Thought prompting benefits multi-step inference but yields mixed results elsewhere. EpiQAL provides fine-grained diagnostic signals for evidence-grounding, inferential reasoning, and conclusion reconstruction.
title EpiQAL: Benchmarking Large Language Models in Epidemiological Question Answering and Reasoning
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2601.03471