TeaRAG: A Token-Efficient Agentic Retrieval-Augmented Generation Framework

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Chao, Wang, Yuhao, Xu, Derong, Zhang, Haoxin, Lyu, Yuanjie, Chen, Yuhao, Liu, Shuochen, Xu, Tong, Zhao, Xiangyu, Gao, Yan, Hu, Yao, Chen, Enhong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918190830845952
author Zhang, Chao
Wang, Yuhao
Xu, Derong
Zhang, Haoxin
Lyu, Yuanjie
Chen, Yuhao
Liu, Shuochen
Xu, Tong
Zhao, Xiangyu
Gao, Yan
Hu, Yao
Chen, Enhong
author_facet Zhang, Chao
Wang, Yuhao
Xu, Derong
Zhang, Haoxin
Lyu, Yuanjie
Chen, Yuhao
Liu, Shuochen
Xu, Tong
Zhao, Xiangyu
Gao, Yan
Hu, Yao
Chen, Enhong
contents Retrieval-Augmented Generation (RAG) utilizes external knowledge to augment Large Language Models' (LLMs) reliability. For flexibility, agentic RAG employs autonomous, multi-round retrieval and reasoning to resolve queries. Although recent agentic RAG has improved via reinforcement learning, they often incur substantial token overhead from search and reasoning processes. This trade-off prioritizes accuracy over efficiency. To address this issue, this work proposes TeaRAG, a token-efficient agentic RAG framework capable of compressing both retrieval content and reasoning steps. 1) First, the retrieved content is compressed by augmenting chunk-based semantic retrieval with a graph retrieval using concise triplets. A knowledge association graph is then built from semantic similarity and co-occurrence. Finally, Personalized PageRank is leveraged to highlight key knowledge within this graph, reducing the number of tokens per retrieval. 2) Besides, to reduce reasoning steps, Iterative Process-aware Direct Preference Optimization (IP-DPO) is proposed. Specifically, our reward function evaluates the knowledge sufficiency by a knowledge matching mechanism, while penalizing excessive reasoning steps. This design can produce high-quality preference-pair datasets, supporting iterative DPO to improve reasoning conciseness. Across six datasets, TeaRAG improves the average Exact Match by 4% and 2% while reducing output tokens by 61% and 59% on Llama3-8B-Instruct and Qwen2.5-14B-Instruct, respectively. Code is available at https://github.com/Applied-Machine-Learning-Lab/TeaRAG.
format Preprint
id arxiv_https___arxiv_org_abs_2511_05385
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TeaRAG: A Token-Efficient Agentic Retrieval-Augmented Generation Framework
Zhang, Chao
Wang, Yuhao
Xu, Derong
Zhang, Haoxin
Lyu, Yuanjie
Chen, Yuhao
Liu, Shuochen
Xu, Tong
Zhao, Xiangyu
Gao, Yan
Hu, Yao
Chen, Enhong
Information Retrieval
Artificial Intelligence
Retrieval-Augmented Generation (RAG) utilizes external knowledge to augment Large Language Models' (LLMs) reliability. For flexibility, agentic RAG employs autonomous, multi-round retrieval and reasoning to resolve queries. Although recent agentic RAG has improved via reinforcement learning, they often incur substantial token overhead from search and reasoning processes. This trade-off prioritizes accuracy over efficiency. To address this issue, this work proposes TeaRAG, a token-efficient agentic RAG framework capable of compressing both retrieval content and reasoning steps. 1) First, the retrieved content is compressed by augmenting chunk-based semantic retrieval with a graph retrieval using concise triplets. A knowledge association graph is then built from semantic similarity and co-occurrence. Finally, Personalized PageRank is leveraged to highlight key knowledge within this graph, reducing the number of tokens per retrieval. 2) Besides, to reduce reasoning steps, Iterative Process-aware Direct Preference Optimization (IP-DPO) is proposed. Specifically, our reward function evaluates the knowledge sufficiency by a knowledge matching mechanism, while penalizing excessive reasoning steps. This design can produce high-quality preference-pair datasets, supporting iterative DPO to improve reasoning conciseness. Across six datasets, TeaRAG improves the average Exact Match by 4% and 2% while reducing output tokens by 61% and 59% on Llama3-8B-Instruct and Qwen2.5-14B-Instruct, respectively. Code is available at https://github.com/Applied-Machine-Learning-Lab/TeaRAG.
title TeaRAG: A Token-Efficient Agentic Retrieval-Augmented Generation Framework
topic Information Retrieval
Artificial Intelligence
url https://arxiv.org/abs/2511.05385