Rethinking the Reranker: Boundary-Aware Evidence Selection for Robust Retrieval-Augmented Generation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Sun, Jiashuo, Jiang, Pengcheng, Wang, Saizhuo, Fan, Jiajun, Wang, Heng, Ouyang, Siru, Zhong, Ming, Jiao, Yizhu, Huang, Chengsong, Xu, Xueqiang, Han, Pengrui, Li, Peiran, Huang, Jiaxin, Liu, Ge, Ji, Heng, Han, Jiawei
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908809727836160
author Sun, Jiashuo
Jiang, Pengcheng
Wang, Saizhuo
Fan, Jiajun
Wang, Heng
Ouyang, Siru
Zhong, Ming
Jiao, Yizhu
Huang, Chengsong
Xu, Xueqiang
Han, Pengrui
Li, Peiran
Huang, Jiaxin
Liu, Ge
Ji, Heng
Han, Jiawei
author_facet Sun, Jiashuo
Jiang, Pengcheng
Wang, Saizhuo
Fan, Jiajun
Wang, Heng
Ouyang, Siru
Zhong, Ming
Jiao, Yizhu
Huang, Chengsong
Xu, Xueqiang
Han, Pengrui
Li, Peiran
Huang, Jiaxin
Liu, Ge
Ji, Heng
Han, Jiawei
contents Retrieval-Augmented Generation (RAG) systems remain brittle under realistic retrieval noise, even when the required evidence appears in the top-K results. A key reason is that retrievers and rerankers optimize solely for relevance, often selecting either trivial, answer-revealing passages or evidence that lacks the critical information required to answer the question, without considering whether the evidence is suitable for the generator. We propose BAR-RAG, which reframes the reranker as a boundary-aware evidence selector that targets the generator's Goldilocks Zone -- evidence that is neither trivially easy nor fundamentally unanswerable for the generator, but is challenging yet sufficient for inference and thus provides the strongest learning signal. BAR-RAG trains the selector with reinforcement learning using generator feedback, and adopts a two-stage pipeline that fine-tunes the generator under the induced evidence distribution to mitigate the distribution mismatch between training and inference. Experiments on knowledge-intensive question answering benchmarks show that BAR-RAG consistently improves end-to-end performance under noisy retrieval, achieving an average gain of 10.3 percent over strong RAG and reranking baselines while substantially improving robustness. Code is publicly avaliable at https://github.com/GasolSun36/BAR-RAG.
format Preprint
id arxiv_https___arxiv_org_abs_2602_03689
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Rethinking the Reranker: Boundary-Aware Evidence Selection for Robust Retrieval-Augmented Generation
Sun, Jiashuo
Jiang, Pengcheng
Wang, Saizhuo
Fan, Jiajun
Wang, Heng
Ouyang, Siru
Zhong, Ming
Jiao, Yizhu
Huang, Chengsong
Xu, Xueqiang
Han, Pengrui
Li, Peiran
Huang, Jiaxin
Liu, Ge
Ji, Heng
Han, Jiawei
Computation and Language
Artificial Intelligence
Retrieval-Augmented Generation (RAG) systems remain brittle under realistic retrieval noise, even when the required evidence appears in the top-K results. A key reason is that retrievers and rerankers optimize solely for relevance, often selecting either trivial, answer-revealing passages or evidence that lacks the critical information required to answer the question, without considering whether the evidence is suitable for the generator. We propose BAR-RAG, which reframes the reranker as a boundary-aware evidence selector that targets the generator's Goldilocks Zone -- evidence that is neither trivially easy nor fundamentally unanswerable for the generator, but is challenging yet sufficient for inference and thus provides the strongest learning signal. BAR-RAG trains the selector with reinforcement learning using generator feedback, and adopts a two-stage pipeline that fine-tunes the generator under the induced evidence distribution to mitigate the distribution mismatch between training and inference. Experiments on knowledge-intensive question answering benchmarks show that BAR-RAG consistently improves end-to-end performance under noisy retrieval, achieving an average gain of 10.3 percent over strong RAG and reranking baselines while substantially improving robustness. Code is publicly avaliable at https://github.com/GasolSun36/BAR-RAG.
title Rethinking the Reranker: Boundary-Aware Evidence Selection for Robust Retrieval-Augmented Generation
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2602.03689