ClueAnchor: Clue-Anchored Knowledge Reasoning Exploration and Optimization for Retrieval-Augmented Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Hao, Yan, Yukun, Mei, Sen, Che, Wanxiang, Liu, Zhenghao, Shi, Qi, Li, Xinze, Fan, Yuchun, Huang, Pengcheng, Xiong, Qiushi, Liu, Zhiyuan, Sun, Maosong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917049771491328
author Chen, Hao
Yan, Yukun
Mei, Sen
Che, Wanxiang
Liu, Zhenghao
Shi, Qi
Li, Xinze
Fan, Yuchun
Huang, Pengcheng
Xiong, Qiushi
Liu, Zhiyuan
Sun, Maosong
author_facet Chen, Hao
Yan, Yukun
Mei, Sen
Che, Wanxiang
Liu, Zhenghao
Shi, Qi
Li, Xinze
Fan, Yuchun
Huang, Pengcheng
Xiong, Qiushi
Liu, Zhiyuan
Sun, Maosong
contents Retrieval-Augmented Generation (RAG) augments Large Language Models (LLMs) with external knowledge to improve factuality. However, existing RAG systems frequently underutilize the retrieved documents, failing to extract and integrate the key clues needed to support faithful and interpretable reasoning, especially in cases where relevant evidence is implicit, scattered, or obscured by noise. To address this issue, we propose ClueAnchor, a novel framework for enhancing RAG via clue-anchored reasoning exploration and optimization. ClueAnchor extracts key clues from retrieved content and generates multiple reasoning paths based on different knowledge configurations, optimizing the model by selecting the most appropriate reasoning path for the given context through reward-based preference optimization. Experiments show that ClueAnchor significantly outperforms prior RAG baselines in the completeness and robustness of reasoning. Further analysis confirms its strong resilience to noisy or partially relevant retrieved content, as well as its capability to identify supporting evidence even in the absence of explicit clue supervision during inference. All codes are available at https://github.com/thunlp/ClueAnchor.
format Preprint
id arxiv_https___arxiv_org_abs_2505_24388
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ClueAnchor: Clue-Anchored Knowledge Reasoning Exploration and Optimization for Retrieval-Augmented Generation
Chen, Hao
Yan, Yukun
Mei, Sen
Che, Wanxiang
Liu, Zhenghao
Shi, Qi
Li, Xinze
Fan, Yuchun
Huang, Pengcheng
Xiong, Qiushi
Liu, Zhiyuan
Sun, Maosong
Computation and Language
Retrieval-Augmented Generation (RAG) augments Large Language Models (LLMs) with external knowledge to improve factuality. However, existing RAG systems frequently underutilize the retrieved documents, failing to extract and integrate the key clues needed to support faithful and interpretable reasoning, especially in cases where relevant evidence is implicit, scattered, or obscured by noise. To address this issue, we propose ClueAnchor, a novel framework for enhancing RAG via clue-anchored reasoning exploration and optimization. ClueAnchor extracts key clues from retrieved content and generates multiple reasoning paths based on different knowledge configurations, optimizing the model by selecting the most appropriate reasoning path for the given context through reward-based preference optimization. Experiments show that ClueAnchor significantly outperforms prior RAG baselines in the completeness and robustness of reasoning. Further analysis confirms its strong resilience to noisy or partially relevant retrieved content, as well as its capability to identify supporting evidence even in the absence of explicit clue supervision during inference. All codes are available at https://github.com/thunlp/ClueAnchor.
title ClueAnchor: Clue-Anchored Knowledge Reasoning Exploration and Optimization for Retrieval-Augmented Generation
topic Computation and Language
url https://arxiv.org/abs/2505.24388