QCFuse: Query-Centric Cache Fusion for Efficient RAG Inference

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Yan, Jianxin, Qian, Zeheng, Ni, Wangze, Shen, Zhitao, Wang, Zhiping, Li, Haoyang, Zhu, Jia, Chen, Lei, Ren, Kui
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914462842224640
author Yan, Jianxin
Qian, Zeheng
Ni, Wangze
Shen, Zhitao
Wang, Zhiping
Li, Haoyang
Zhu, Jia
Chen, Lei
Ren, Kui
author_facet Yan, Jianxin
Qian, Zeheng
Ni, Wangze
Shen, Zhitao
Wang, Zhiping
Li, Haoyang
Zhu, Jia
Chen, Lei
Ren, Kui
contents Cache fusion accelerates generation process of LLMs equipped with RAG through KV caching and selective token recomputation, thereby reducing computational costs and improving efficiency. However, existing methods primarily rely on local perspectives for token selection and lack global awareness from the user query. Utilizing this global awareness is challenging due to the high cost of obtaining context-aware query representations and the strict pipeline constraints required for efficient attention analysis. Thus, this demonstration introduces QCFuse, an innovative KV cache fusion system centered on the user query. QCFuse leverages semantic summary anchors to enhance query representations and selectively recomputes query-related tokens to improve accuracy, updating tokens based on the attention distribution of the most critical Transformer layer to preserve the high efficiency of the pipeline structure. Evaluations on real-world datasets demonstrate that QCFuse significantly improves the response efficiency of LLMs by 40\% while maintaining equivalent accuracy compared to current methods. Additionally, in certain scenarios, QCFuse achieves an attention denoising effect that yields higher response accuracy, demonstrating substantial potential in the optimization of LLM inference.
format Preprint
id arxiv_https___arxiv_org_abs_2604_08585
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle QCFuse: Query-Centric Cache Fusion for Efficient RAG Inference
Yan, Jianxin
Qian, Zeheng
Ni, Wangze
Shen, Zhitao
Wang, Zhiping
Li, Haoyang
Zhu, Jia
Chen, Lei
Ren, Kui
Databases
Artificial Intelligence
Cache fusion accelerates generation process of LLMs equipped with RAG through KV caching and selective token recomputation, thereby reducing computational costs and improving efficiency. However, existing methods primarily rely on local perspectives for token selection and lack global awareness from the user query. Utilizing this global awareness is challenging due to the high cost of obtaining context-aware query representations and the strict pipeline constraints required for efficient attention analysis. Thus, this demonstration introduces QCFuse, an innovative KV cache fusion system centered on the user query. QCFuse leverages semantic summary anchors to enhance query representations and selectively recomputes query-related tokens to improve accuracy, updating tokens based on the attention distribution of the most critical Transformer layer to preserve the high efficiency of the pipeline structure. Evaluations on real-world datasets demonstrate that QCFuse significantly improves the response efficiency of LLMs by 40\% while maintaining equivalent accuracy compared to current methods. Additionally, in certain scenarios, QCFuse achieves an attention denoising effect that yields higher response accuracy, demonstrating substantial potential in the optimization of LLM inference.
title QCFuse: Query-Centric Cache Fusion for Efficient RAG Inference
topic Databases
Artificial Intelligence
url https://arxiv.org/abs/2604.08585