SciEGQA: A Dataset for Scientific Evidence-Grounded Question Answering and Reasoning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yu, Wenhan, Zhang, Zhaoxi, Chen, Wang, Qi, Guanqiang, Li, Weikang, Sha, Lei, Xia, Deguo, Huang, Jizhou
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915897830014976
author Yu, Wenhan
Zhang, Zhaoxi
Chen, Wang
Qi, Guanqiang
Li, Weikang
Sha, Lei
Xia, Deguo
Huang, Jizhou
author_facet Yu, Wenhan
Zhang, Zhaoxi
Chen, Wang
Qi, Guanqiang
Li, Weikang
Sha, Lei
Xia, Deguo
Huang, Jizhou
contents Scientific documents contain complex multimodal structures, which makes evidence localization and scientific reasoning in Document Visual Question Answering particularly challenging. However, most existing benchmarks evaluate models only at the page level without explicitly annotating the evidence regions that support the answer, which limits both interpretability and the reliability of evaluation. To address this limitation, we introduce SciEGQA, a scientific document question answering and reasoning dataset with semantic evidence grounding, where supporting evidence is represented as semantically coherent document regions annotated with bounding boxes. SciEGQA consists of two components: a **human-annotated fine-grained benchmark** containing 1,623 high-quality question--answer pairs, and a **large-scale automatically constructed training set** with over 30K QA pairs generated through an automated data construction pipeline. Extensive experiments on a wide range of Vision-Language Models (VLMs) show that existing models still struggle with evidence localization and evidence-based question answering in scientific documents. Training on the proposed dataset significantly improves the scientific reasoning capabilities of VLMs. The project page is available at https://yuwenhan07.github.io/SciEGQA-project/.
format Preprint
id arxiv_https___arxiv_org_abs_2511_15090
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SciEGQA: A Dataset for Scientific Evidence-Grounded Question Answering and Reasoning
Yu, Wenhan
Zhang, Zhaoxi
Chen, Wang
Qi, Guanqiang
Li, Weikang
Sha, Lei
Xia, Deguo
Huang, Jizhou
Databases
Artificial Intelligence
Computer Vision and Pattern Recognition
Scientific documents contain complex multimodal structures, which makes evidence localization and scientific reasoning in Document Visual Question Answering particularly challenging. However, most existing benchmarks evaluate models only at the page level without explicitly annotating the evidence regions that support the answer, which limits both interpretability and the reliability of evaluation. To address this limitation, we introduce SciEGQA, a scientific document question answering and reasoning dataset with semantic evidence grounding, where supporting evidence is represented as semantically coherent document regions annotated with bounding boxes. SciEGQA consists of two components: a **human-annotated fine-grained benchmark** containing 1,623 high-quality question--answer pairs, and a **large-scale automatically constructed training set** with over 30K QA pairs generated through an automated data construction pipeline. Extensive experiments on a wide range of Vision-Language Models (VLMs) show that existing models still struggle with evidence localization and evidence-based question answering in scientific documents. Training on the proposed dataset significantly improves the scientific reasoning capabilities of VLMs. The project page is available at https://yuwenhan07.github.io/SciEGQA-project/.
title SciEGQA: A Dataset for Scientific Evidence-Grounded Question Answering and Reasoning
topic Databases
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.15090