DeepEra: A Deep Evidence Reranking Agent for Scientific Retrieval-Augmented Generated Question Answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Haotian, Long, Qingqing, Pu, Siyu, Luo, Xiao, Ju, Wei, Xiao, Meng, Zhou, Yuanchun, Zhao, Jianghua, Wang, Xuezhi
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915749836095488
author Chen, Haotian
Long, Qingqing
Pu, Siyu
Luo, Xiao
Ju, Wei
Xiao, Meng
Zhou, Yuanchun
Zhao, Jianghua
Wang, Xuezhi
author_facet Chen, Haotian
Long, Qingqing
Pu, Siyu
Luo, Xiao
Ju, Wei
Xiao, Meng
Zhou, Yuanchun
Zhao, Jianghua
Wang, Xuezhi
contents With the rapid growth of scientific literature, scientific question answering (SciQA) has become increasingly critical for exploring and utilizing scientific knowledge. Retrieval-Augmented Generation (RAG) enhances LLMs by incorporating knowledge from external sources, thereby providing credible evidence for scientific question answering. But existing retrieval and reranking methods remain vulnerable to passages that are semantically similar but logically irrelevant, often reducing factual reliability and amplifying hallucinations.To address this challenge, we propose a Deep Evidence Reranking Agent (DeepEra) that integrates step-by-step reasoning, enabling more precise evaluation of candidate passages beyond surface-level semantics. To support systematic evaluation, we construct SciRAG-SSLI (Scientific RAG - Semantically Similar but Logically Irrelevant), a large-scale dataset comprising about 300K SciQA instances across 10 subjects, constructed from 10M scientific corpus. The dataset combines naturally retrieved contexts with systematically generated distractors to test logical robustness and factual grounding. Comprehensive evaluations confirm that our approach achieves superior retrieval performance compared to leading rerankers. To our knowledge, this work is the first to comprehensively study and empirically validate innegligible SSLI issues in two-stage RAG frameworks.
format Preprint
id arxiv_https___arxiv_org_abs_2601_16478
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DeepEra: A Deep Evidence Reranking Agent for Scientific Retrieval-Augmented Generated Question Answering
Chen, Haotian
Long, Qingqing
Pu, Siyu
Luo, Xiao
Ju, Wei
Xiao, Meng
Zhou, Yuanchun
Zhao, Jianghua
Wang, Xuezhi
Computation and Language
Artificial Intelligence
With the rapid growth of scientific literature, scientific question answering (SciQA) has become increasingly critical for exploring and utilizing scientific knowledge. Retrieval-Augmented Generation (RAG) enhances LLMs by incorporating knowledge from external sources, thereby providing credible evidence for scientific question answering. But existing retrieval and reranking methods remain vulnerable to passages that are semantically similar but logically irrelevant, often reducing factual reliability and amplifying hallucinations.To address this challenge, we propose a Deep Evidence Reranking Agent (DeepEra) that integrates step-by-step reasoning, enabling more precise evaluation of candidate passages beyond surface-level semantics. To support systematic evaluation, we construct SciRAG-SSLI (Scientific RAG - Semantically Similar but Logically Irrelevant), a large-scale dataset comprising about 300K SciQA instances across 10 subjects, constructed from 10M scientific corpus. The dataset combines naturally retrieved contexts with systematically generated distractors to test logical robustness and factual grounding. Comprehensive evaluations confirm that our approach achieves superior retrieval performance compared to leading rerankers. To our knowledge, this work is the first to comprehensively study and empirically validate innegligible SSLI issues in two-stage RAG frameworks.
title DeepEra: A Deep Evidence Reranking Agent for Scientific Retrieval-Augmented Generated Question Answering
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2601.16478