Rethinking the Reranker: Boundary-Aware Evidence Selection for Robust Retrieval-Augmented Generation
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , , , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866908809727836160 |
|---|---|
| author | Sun, Jiashuo Jiang, Pengcheng Wang, Saizhuo Fan, Jiajun Wang, Heng Ouyang, Siru Zhong, Ming Jiao, Yizhu Huang, Chengsong Xu, Xueqiang Han, Pengrui Li, Peiran Huang, Jiaxin Liu, Ge Ji, Heng Han, Jiawei |
| author_facet | Sun, Jiashuo Jiang, Pengcheng Wang, Saizhuo Fan, Jiajun Wang, Heng Ouyang, Siru Zhong, Ming Jiao, Yizhu Huang, Chengsong Xu, Xueqiang Han, Pengrui Li, Peiran Huang, Jiaxin Liu, Ge Ji, Heng Han, Jiawei |
| contents | Retrieval-Augmented Generation (RAG) systems remain brittle under realistic retrieval noise, even when the required evidence appears in the top-K results. A key reason is that retrievers and rerankers optimize solely for relevance, often selecting either trivial, answer-revealing passages or evidence that lacks the critical information required to answer the question, without considering whether the evidence is suitable for the generator. We propose BAR-RAG, which reframes the reranker as a boundary-aware evidence selector that targets the generator's Goldilocks Zone -- evidence that is neither trivially easy nor fundamentally unanswerable for the generator, but is challenging yet sufficient for inference and thus provides the strongest learning signal. BAR-RAG trains the selector with reinforcement learning using generator feedback, and adopts a two-stage pipeline that fine-tunes the generator under the induced evidence distribution to mitigate the distribution mismatch between training and inference. Experiments on knowledge-intensive question answering benchmarks show that BAR-RAG consistently improves end-to-end performance under noisy retrieval, achieving an average gain of 10.3 percent over strong RAG and reranking baselines while substantially improving robustness. Code is publicly avaliable at https://github.com/GasolSun36/BAR-RAG. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_03689 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Rethinking the Reranker: Boundary-Aware Evidence Selection for Robust Retrieval-Augmented Generation Sun, Jiashuo Jiang, Pengcheng Wang, Saizhuo Fan, Jiajun Wang, Heng Ouyang, Siru Zhong, Ming Jiao, Yizhu Huang, Chengsong Xu, Xueqiang Han, Pengrui Li, Peiran Huang, Jiaxin Liu, Ge Ji, Heng Han, Jiawei Computation and Language Artificial Intelligence Retrieval-Augmented Generation (RAG) systems remain brittle under realistic retrieval noise, even when the required evidence appears in the top-K results. A key reason is that retrievers and rerankers optimize solely for relevance, often selecting either trivial, answer-revealing passages or evidence that lacks the critical information required to answer the question, without considering whether the evidence is suitable for the generator. We propose BAR-RAG, which reframes the reranker as a boundary-aware evidence selector that targets the generator's Goldilocks Zone -- evidence that is neither trivially easy nor fundamentally unanswerable for the generator, but is challenging yet sufficient for inference and thus provides the strongest learning signal. BAR-RAG trains the selector with reinforcement learning using generator feedback, and adopts a two-stage pipeline that fine-tunes the generator under the induced evidence distribution to mitigate the distribution mismatch between training and inference. Experiments on knowledge-intensive question answering benchmarks show that BAR-RAG consistently improves end-to-end performance under noisy retrieval, achieving an average gain of 10.3 percent over strong RAG and reranking baselines while substantially improving robustness. Code is publicly avaliable at https://github.com/GasolSun36/BAR-RAG. |
| title | Rethinking the Reranker: Boundary-Aware Evidence Selection for Robust Retrieval-Augmented Generation |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2602.03689 |