HyperRAG: Enhancing Quality-Efficiency Tradeoffs in Retrieval-Augmented Generation with Reranker KV-Cache Reuse
Fuente:
arXiv
Salvato in:
| Autori principali: | An, Yuwei, Cheng, Yihua, Park, Seo Jin, Jiang, Junchen |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
HyperRAG: Reasoning N-ary Facts over Hypergraphs for Retrieval Augmented Generation
di: Lien, Wen-Sheng, et al.
Pubblicazione: (2026)
di: Lien, Wen-Sheng, et al.
Pubblicazione: (2026)
Towards More Economical Context-Augmented LLM Generation by Reusing Stored KV Cache
di: Li, Hanchen, et al.
Pubblicazione: (2025)
di: Li, Hanchen, et al.
Pubblicazione: (2025)
TurboRAG: Accelerating Retrieval-Augmented Generation with Precomputed KV Caches for Chunked Text
di: Lu, Songshuo, et al.
Pubblicazione: (2024)
di: Lu, Songshuo, et al.
Pubblicazione: (2024)
Enhancing Financial Report Question-Answering: A Retrieval-Augmented Generation System with Reranking Analysis
di: Cheng, Zhiyuan, et al.
Pubblicazione: (2026)
di: Cheng, Zhiyuan, et al.
Pubblicazione: (2026)
From Prefix Cache to Fusion RAG Cache: Accelerating LLM Inference in Retrieval-Augmented Generation
di: Wang, Jiahao, et al.
Pubblicazione: (2026)
di: Wang, Jiahao, et al.
Pubblicazione: (2026)
Inference-Time Hyper-Scaling with KV Cache Compression
di: Łańcucki, Adrian, et al.
Pubblicazione: (2025)
di: Łańcucki, Adrian, et al.
Pubblicazione: (2025)
Speculative RAG: Enhancing Retrieval Augmented Generation through Drafting
di: Wang, Zilong, et al.
Pubblicazione: (2024)
di: Wang, Zilong, et al.
Pubblicazione: (2024)
QAQ: Quality Adaptive Quantization for LLM KV Cache
di: Dong, Shichen, et al.
Pubblicazione: (2024)
di: Dong, Shichen, et al.
Pubblicazione: (2024)
EFIM: Efficient Serving of LLMs for Infilling Tasks with Improved KV Cache Reuse
di: Guo, Tianyu, et al.
Pubblicazione: (2025)
di: Guo, Tianyu, et al.
Pubblicazione: (2025)
KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse
di: Yang, Jingbo, et al.
Pubblicazione: (2025)
di: Yang, Jingbo, et al.
Pubblicazione: (2025)
DynamicRAG: Leveraging Outputs of Large Language Model as Feedback for Dynamic Reranking in Retrieval-Augmented Generation
di: Sun, Jiashuo, et al.
Pubblicazione: (2025)
di: Sun, Jiashuo, et al.
Pubblicazione: (2025)
DroidSpeak: KV Cache Sharing for Cross-LLM Communication and Multi-LLM Serving
di: Liu, Yuhan, et al.
Pubblicazione: (2024)
di: Liu, Yuhan, et al.
Pubblicazione: (2024)
RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation
di: Jin, Chao, et al.
Pubblicazione: (2024)
di: Jin, Chao, et al.
Pubblicazione: (2024)
AutoRAG-HP: Automatic Online Hyper-Parameter Tuning for Retrieval-Augmented Generation
di: Fu, Jia, et al.
Pubblicazione: (2024)
di: Fu, Jia, et al.
Pubblicazione: (2024)
KVSculpt: KV Cache Compression as Distillation
di: Jiang, Bo, et al.
Pubblicazione: (2026)
di: Jiang, Bo, et al.
Pubblicazione: (2026)
LongRAG: Enhancing Retrieval-Augmented Generation with Long-context LLMs
di: Jiang, Ziyan, et al.
Pubblicazione: (2024)
di: Jiang, Ziyan, et al.
Pubblicazione: (2024)
Grounded Cache Routing for Retrieval-Augmented Generation: When Is It Safe to Reuse an Answer?
di: Shah, Syed Huma
Pubblicazione: (2026)
di: Shah, Syed Huma
Pubblicazione: (2026)
SciRerankBench: Benchmarking Rerankers Towards Scientific Retrieval-Augmented Generated LLMs
di: Chen, Haotian, et al.
Pubblicazione: (2025)
di: Chen, Haotian, et al.
Pubblicazione: (2025)
CacheFocus: Dynamic Cache Re-Positioning for Efficient Retrieval-Augmented Generation
di: Lee, Kun-Hui, et al.
Pubblicazione: (2025)
di: Lee, Kun-Hui, et al.
Pubblicazione: (2025)
TrustRAG: Enhancing Robustness and Trustworthiness in Retrieval-Augmented Generation
di: Zhou, Huichi, et al.
Pubblicazione: (2025)
di: Zhou, Huichi, et al.
Pubblicazione: (2025)
KiRAG: Knowledge-Driven Iterative Retriever for Enhancing Retrieval-Augmented Generation
di: Fang, Jinyuan, et al.
Pubblicazione: (2025)
di: Fang, Jinyuan, et al.
Pubblicazione: (2025)
ParisKV: Fast and Drift-Robust KV-Cache Retrieval for Long-Context LLMs
di: Qi, Yanlin, et al.
Pubblicazione: (2026)
di: Qi, Yanlin, et al.
Pubblicazione: (2026)
KVShare: An LLM Service System with Efficient and Effective Multi-Tenant KV Cache Reuse
di: Yang, Huan, et al.
Pubblicazione: (2025)
di: Yang, Huan, et al.
Pubblicazione: (2025)
InfoGain-RAG: Boosting Retrieval-Augmented Generation via Document Information Gain-based Reranking and Filtering
di: Wang, Zihan, et al.
Pubblicazione: (2025)
di: Wang, Zihan, et al.
Pubblicazione: (2025)
ValuesRAG: Enhancing Cultural Alignment Through Retrieval-Augmented Contextual Learning
di: Seo, Wonduk, et al.
Pubblicazione: (2025)
di: Seo, Wonduk, et al.
Pubblicazione: (2025)
Do Large Language Models Need a Content Delivery Network?
di: Cheng, Yihua, et al.
Pubblicazione: (2024)
di: Cheng, Yihua, et al.
Pubblicazione: (2024)
MobileRAG: Enhancing Mobile Agent with Retrieval-Augmented Generation
di: Loo, Gowen, et al.
Pubblicazione: (2025)
di: Loo, Gowen, et al.
Pubblicazione: (2025)
M$^3$KG-RAG: Multi-hop Multimodal Knowledge Graph-enhanced Retrieval-Augmented Generation
di: Park, Hyeongcheol, et al.
Pubblicazione: (2025)
di: Park, Hyeongcheol, et al.
Pubblicazione: (2025)
LLM-Confidence Reranker: A Training-Free Approach for Enhancing Retrieval-Augmented Generation Systems
di: Song, Zhipeng, et al.
Pubblicazione: (2026)
di: Song, Zhipeng, et al.
Pubblicazione: (2026)
HeteroCache: A Dynamic Retrieval Approach to Heterogeneous KV Cache Compression for Long-Context LLM Inference
di: Shi, Zhiyuan, et al.
Pubblicazione: (2026)
di: Shi, Zhiyuan, et al.
Pubblicazione: (2026)
HyperbolicRAG: Enhancing Retrieval-Augmented Generation with Hyperbolic Representations
di: Cao, Linxiao, et al.
Pubblicazione: (2025)
di: Cao, Linxiao, et al.
Pubblicazione: (2025)
RAG+: Enhancing Retrieval-Augmented Generation with Application-Aware Reasoning
di: Wang, Yu, et al.
Pubblicazione: (2025)
di: Wang, Yu, et al.
Pubblicazione: (2025)
FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference
di: Liu, Guangda, et al.
Pubblicazione: (2025)
di: Liu, Guangda, et al.
Pubblicazione: (2025)
Rethinking the Reranker: Boundary-Aware Evidence Selection for Robust Retrieval-Augmented Generation
di: Sun, Jiashuo, et al.
Pubblicazione: (2026)
di: Sun, Jiashuo, et al.
Pubblicazione: (2026)
Don't Do RAG: When Cache-Augmented Generation is All You Need for Knowledge Tasks
di: Chan, Brian J, et al.
Pubblicazione: (2024)
di: Chan, Brian J, et al.
Pubblicazione: (2024)
Enhancing Retrieval and Managing Retrieval: A Four-Module Synergy for Improved Quality and Efficiency in RAG Systems
di: Shi, Yunxiao, et al.
Pubblicazione: (2024)
di: Shi, Yunxiao, et al.
Pubblicazione: (2024)
CAR: Query-Guided Confidence-Aware Reranking for Retrieval-Augmented Generation
di: Song, Zhipeng, et al.
Pubblicazione: (2026)
di: Song, Zhipeng, et al.
Pubblicazione: (2026)
Zero-RAG: Towards Retrieval-Augmented Generation with Zero Redundant Knowledge
di: Luo, Qi, et al.
Pubblicazione: (2025)
di: Luo, Qi, et al.
Pubblicazione: (2025)
GraphKV: Breaking the Static Selection Paradigm with Graph-Based KV Cache Eviction
di: Li, Xuelin, et al.
Pubblicazione: (2025)
di: Li, Xuelin, et al.
Pubblicazione: (2025)
MCTS-RAG: Enhancing Retrieval-Augmented Generation with Monte Carlo Tree Search
di: Hu, Yunhai, et al.
Pubblicazione: (2025)
di: Hu, Yunhai, et al.
Pubblicazione: (2025)
Documenti analoghi
-
HyperRAG: Reasoning N-ary Facts over Hypergraphs for Retrieval Augmented Generation
di: Lien, Wen-Sheng, et al.
Pubblicazione: (2026) -
Towards More Economical Context-Augmented LLM Generation by Reusing Stored KV Cache
di: Li, Hanchen, et al.
Pubblicazione: (2025) -
TurboRAG: Accelerating Retrieval-Augmented Generation with Precomputed KV Caches for Chunked Text
di: Lu, Songshuo, et al.
Pubblicazione: (2024) -
Enhancing Financial Report Question-Answering: A Retrieval-Augmented Generation System with Reranking Analysis
di: Cheng, Zhiyuan, et al.
Pubblicazione: (2026) -
From Prefix Cache to Fusion RAG Cache: Accelerating LLM Inference in Retrieval-Augmented Generation
di: Wang, Jiahao, et al.
Pubblicazione: (2026)