RARE: Retrieval-Augmented Reasoning Modeling

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Zhengren, Yu, Jiayang, Ma, Dongsheng, Chen, Zhe, Wang, Yu, Li, Zhiyu, Xiong, Feiyu, Wang, Yanfeng, E, Weinan, Tang, Linpeng, Zhang, Wentao
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910949780226048
author Wang, Zhengren
Yu, Jiayang
Ma, Dongsheng
Chen, Zhe
Wang, Yu
Li, Zhiyu
Xiong, Feiyu
Wang, Yanfeng
E, Weinan
Tang, Linpeng
Zhang, Wentao
author_facet Wang, Zhengren
Yu, Jiayang
Ma, Dongsheng
Chen, Zhe
Wang, Yu
Li, Zhiyu
Xiong, Feiyu
Wang, Yanfeng
E, Weinan
Tang, Linpeng
Zhang, Wentao
contents Domain-specific intelligence demands specialized knowledge and sophisticated reasoning for problem-solving, posing significant challenges for large language models (LLMs) that struggle with knowledge hallucination and inadequate reasoning capabilities under constrained parameter budgets. Inspired by Bloom's Taxonomy in educational theory, we propose Retrieval-Augmented Reasoning Modeling (RARE), a novel paradigm that decouples knowledge storage from reasoning optimization. RARE externalizes domain knowledge to retrievable sources and internalizes domain-specific reasoning patterns during training. Specifically, by injecting retrieved knowledge into training prompts with masked losses, RARE transforms learning objectives from rote memorization to contextualized reasoning. It enables models to bypass parameter-intensive memorization and prioritize the development of higher-order cognitive processes. Extensive experiments demonstrate that lightweight RARE-trained models (e.g., Llama-3.1-8B) could achieve state-of-the-art performance, surpassing retrieval-augmented GPT-4 and DeepSeek-R1 up to approximately 20\% accuracy. RARE establishes a paradigm shift where maintainable external knowledge bases synergize with compact, reasoning-optimized models, collectively driving more scalable domain-specific intelligence.
format Preprint
id arxiv_https___arxiv_org_abs_2503_23513
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RARE: Retrieval-Augmented Reasoning Modeling
Wang, Zhengren
Yu, Jiayang
Ma, Dongsheng
Chen, Zhe
Wang, Yu
Li, Zhiyu
Xiong, Feiyu
Wang, Yanfeng
E, Weinan
Tang, Linpeng
Zhang, Wentao
Computation and Language
Domain-specific intelligence demands specialized knowledge and sophisticated reasoning for problem-solving, posing significant challenges for large language models (LLMs) that struggle with knowledge hallucination and inadequate reasoning capabilities under constrained parameter budgets. Inspired by Bloom's Taxonomy in educational theory, we propose Retrieval-Augmented Reasoning Modeling (RARE), a novel paradigm that decouples knowledge storage from reasoning optimization. RARE externalizes domain knowledge to retrievable sources and internalizes domain-specific reasoning patterns during training. Specifically, by injecting retrieved knowledge into training prompts with masked losses, RARE transforms learning objectives from rote memorization to contextualized reasoning. It enables models to bypass parameter-intensive memorization and prioritize the development of higher-order cognitive processes. Extensive experiments demonstrate that lightweight RARE-trained models (e.g., Llama-3.1-8B) could achieve state-of-the-art performance, surpassing retrieval-augmented GPT-4 and DeepSeek-R1 up to approximately 20\% accuracy. RARE establishes a paradigm shift where maintainable external knowledge bases synergize with compact, reasoning-optimized models, collectively driving more scalable domain-specific intelligence.
title RARE: Retrieval-Augmented Reasoning Modeling
topic Computation and Language
url https://arxiv.org/abs/2503.23513