Accelerating Inference of Retrieval-Augmented Generation via Sparse Context Selection
Fuente:
arXiv
Saved in:
| Main Authors: | Zhu, Yun, Gu, Jia-Chen, Sikora, Caitlin, Ko, Ho, Liu, Yinxiao, Lin, Chu-Cheng, Shu, Lei, Luo, Liangchen, Meng, Lei, Liu, Bang, Chen, Jindong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Fusion-Eval: Integrating Assistant Evaluators with LLMs
by: Shu, Lei, et al.
Published: (2023)
by: Shu, Lei, et al.
Published: (2023)
Beyond Sparse Rewards: Enhancing Reinforcement Learning with Language Model Critique in Text Generation
by: Cao, Meng, et al.
Published: (2024)
by: Cao, Meng, et al.
Published: (2024)
Improve Mathematical Reasoning in Language Models by Automated Process Supervision
by: Luo, Liangchen, et al.
Published: (2024)
by: Luo, Liangchen, et al.
Published: (2024)
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
by: Liu, Di, et al.
Published: (2024)
by: Liu, Di, et al.
Published: (2024)
Can Knowledge-Graph-based Retrieval Augmented Generation Really Retrieve What You Need?
by: Yu, Junchi, et al.
Published: (2025)
by: Yu, Junchi, et al.
Published: (2025)
Corrective Retrieval Augmented Generation
by: Yan, Shi-Qi, et al.
Published: (2024)
by: Yan, Shi-Qi, et al.
Published: (2024)
VecAttention: Vector-wise Sparse Attention for Accelerating Long Context Inference
by: Liu, Anmin, et al.
Published: (2026)
by: Liu, Anmin, et al.
Published: (2026)
RAPID: Long-Context Inference with Retrieval-Augmented Speculative Decoding
by: Chen, Guanzheng, et al.
Published: (2025)
by: Chen, Guanzheng, et al.
Published: (2025)
SkillRAE: Agent Skill-Based Context Compilation for Retrieval-Augmented Execution
by: Meng, Xiangcheng, et al.
Published: (2026)
by: Meng, Xiangcheng, et al.
Published: (2026)
Influence Guided Context Selection for Effective Retrieval-Augmented Generation
by: Deng, Jiale, et al.
Published: (2025)
by: Deng, Jiale, et al.
Published: (2025)
LycheeDecode: Accelerating Long-Context LLM Inference via Hybrid-Head Sparse Decoding
by: Lin, Gang, et al.
Published: (2026)
by: Lin, Gang, et al.
Published: (2026)
SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention
by: Zhu, Qianchao, et al.
Published: (2024)
by: Zhu, Qianchao, et al.
Published: (2024)
Accelerating Sparse Transformer Inference on GPU
by: Dai, Wenhao, et al.
Published: (2025)
by: Dai, Wenhao, et al.
Published: (2025)
Towards Adaptive Memory-Based Optimization for Enhanced Retrieval-Augmented Generation
by: Qin, Qitao, et al.
Published: (2025)
by: Qin, Qitao, et al.
Published: (2025)
Improving Context Fidelity via Native Retrieval-Augmented Reasoning
by: Wang, Suyuchen, et al.
Published: (2025)
by: Wang, Suyuchen, et al.
Published: (2025)
FlashBack:Efficient Retrieval-Augmented Language Modeling for Long Context Inference
by: Liu, Runheng, et al.
Published: (2024)
by: Liu, Runheng, et al.
Published: (2024)
SEReDeEP: Hallucination Detection in Retrieval-Augmented Models via Semantic Entropy and Context-Parameter Fusion
by: Wang, Lei
Published: (2025)
by: Wang, Lei
Published: (2025)
Tool-MCoT: Tool Augmented Multimodal Chain-of-Thought for Content Safety Moderation
by: Zhang, Shutong, et al.
Published: (2026)
by: Zhang, Shutong, et al.
Published: (2026)
AB-Sparse: Sparse Attention with Adaptive Block Size for Accurate and Efficient Long-Context Inference
by: Liu, Di, et al.
Published: (2026)
by: Liu, Di, et al.
Published: (2026)
WiSparse: Boosting LLM Inference Efficiency with Weight-Aware Mixed Activation Sparsity
by: Chen, Lei, et al.
Published: (2026)
by: Chen, Lei, et al.
Published: (2026)
OpenRAG: Optimizing RAG End-to-End via In-Context Retrieval Learning
by: Zhou, Jiawei, et al.
Published: (2025)
by: Zhou, Jiawei, et al.
Published: (2025)
Inference Scaling for Long-Context Retrieval Augmented Generation
by: Yue, Zhenrui, et al.
Published: (2024)
by: Yue, Zhenrui, et al.
Published: (2024)
Hearing More with Less: Multi-Modal Retrieval-and-Selection Augmented Conversational LLM-Based ASR
by: Mu, Bingshen, et al.
Published: (2025)
by: Mu, Bingshen, et al.
Published: (2025)
IntPro: A Proxy Agent for Context-Aware Intent Understanding via Retrieval-conditioned Inference
by: Liu, Guanming, et al.
Published: (2026)
by: Liu, Guanming, et al.
Published: (2026)
ZigzagAttention: Efficient Long-Context Inference with Exclusive Retrieval and Streaming Heads
by: Liu, Zhuorui, et al.
Published: (2025)
by: Liu, Zhuorui, et al.
Published: (2025)
R^2AG: Incorporating Retrieval Information into Retrieval Augmented Generation
by: Ye, Fuda, et al.
Published: (2024)
by: Ye, Fuda, et al.
Published: (2024)
Generating Is Believing: Membership Inference Attacks against Retrieval-Augmented Generation
by: Li, Yuying, et al.
Published: (2024)
by: Li, Yuying, et al.
Published: (2024)
SparseAccelerate: Efficient Long-Context Inference for Mid-Range GPUs
by: Vo, James
Published: (2024)
by: Vo, James
Published: (2024)
Unlocking Multi-View Insights in Knowledge-Dense Retrieval-Augmented Generation
by: Chen, Guanhua, et al.
Published: (2024)
by: Chen, Guanhua, et al.
Published: (2024)
RASST: Fast Cross-modal Retrieval-Augmented Simultaneous Speech Translation
by: Luo, Jiaxuan, et al.
Published: (2026)
by: Luo, Jiaxuan, et al.
Published: (2026)
FACT: Examining the Effectiveness of Iterative Context Rewriting for Multi-fact Retrieval
by: Wang, Jinlin, et al.
Published: (2024)
by: Wang, Jinlin, et al.
Published: (2024)
From Prefix Cache to Fusion RAG Cache: Accelerating LLM Inference in Retrieval-Augmented Generation
by: Wang, Jiahao, et al.
Published: (2026)
by: Wang, Jiahao, et al.
Published: (2026)
MISA: Mixture of Indexer Sparse Attention for Long-Context LLM Inference
by: Zhou, Ruijie, et al.
Published: (2026)
by: Zhou, Ruijie, et al.
Published: (2026)
Tensor network algorithm to solve polaron impurity problems
by: Chen, Ruofan, et al.
Published: (2025)
by: Chen, Ruofan, et al.
Published: (2025)
True Multimodal In-Context Learning Needs Attention to the Visual Context
by: Chen, Shuo, et al.
Published: (2025)
by: Chen, Shuo, et al.
Published: (2025)
Context Selection and Rewriting for Video-based Educational Question Generation
by: Yu, Mengxia, et al.
Published: (2025)
by: Yu, Mengxia, et al.
Published: (2025)
CoreInfer: Accelerating Large Language Model Inference with Semantics-Inspired Adaptive Sparse Activation
by: Wang, Qinsi, et al.
Published: (2024)
by: Wang, Qinsi, et al.
Published: (2024)
Momentum-Accelerated Richardson(m) and Their Multilevel Neural Solvers
by: Wang, Zhen, et al.
Published: (2024)
by: Wang, Zhen, et al.
Published: (2024)
Towards Automated Smart Contract Generation: Evaluation, Benchmarking, and Retrieval-Augmented Repair
by: Chen, Zaoyu, et al.
Published: (2025)
by: Chen, Zaoyu, et al.
Published: (2025)
TeleRAG: Efficient Retrieval-Augmented Generation Inference with Lookahead Retrieval
by: Lin, Chien-Yu, et al.
Published: (2025)
by: Lin, Chien-Yu, et al.
Published: (2025)
Similar Items
-
Fusion-Eval: Integrating Assistant Evaluators with LLMs
by: Shu, Lei, et al.
Published: (2023) -
Beyond Sparse Rewards: Enhancing Reinforcement Learning with Language Model Critique in Text Generation
by: Cao, Meng, et al.
Published: (2024) -
Improve Mathematical Reasoning in Language Models by Automated Process Supervision
by: Luo, Liangchen, et al.
Published: (2024) -
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
by: Liu, Di, et al.
Published: (2024) -
Can Knowledge-Graph-based Retrieval Augmented Generation Really Retrieve What You Need?
by: Yu, Junchi, et al.
Published: (2025)