TurboRAG: Accelerating Retrieval-Augmented Generation with Precomputed KV Caches for Chunked Text
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Lu, Songshuo, Wang, Hua, Rong, Yutian, Chen, Zhi, Tang, Yaohua |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
URPO: A Unified Reward & Policy Optimization Framework for Large Language Models
par: Lu, Songshuo, et autres
Publié: (2025)
par: Lu, Songshuo, et autres
Publié: (2025)
Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries
par: Wu, Yin, et autres
Publié: (2025)
par: Wu, Yin, et autres
Publié: (2025)
MobileRAG: Enhancing Mobile Agent with Retrieval-Augmented Generation
par: Loo, Gowen, et autres
Publié: (2025)
par: Loo, Gowen, et autres
Publié: (2025)
LaB-RAG: Label Boosted Retrieval Augmented Generation for Radiology Report Generation
par: Song, Steven, et autres
Publié: (2024)
par: Song, Steven, et autres
Publié: (2024)
Cross-modal RAG: Sub-dimensional Text-to-Image Retrieval-Augmented Generation
par: Zhu, Mengdan, et autres
Publié: (2025)
par: Zhu, Mengdan, et autres
Publié: (2025)
StableGS: A Floater-Free Framework for 3D Gaussian Splatting
par: Wang, Luchao, et autres
Publié: (2025)
par: Wang, Luchao, et autres
Publié: (2025)
FastOCR: Dynamic Visual Fixation via KV Cache Pruning for Efficient Document Parsing
par: Tang, Zihan, et autres
Publié: (2026)
par: Tang, Zihan, et autres
Publié: (2026)
VimRAG: Navigating Massive Visual Context in Retrieval-Augmented Generation via Multimodal Memory Graph
par: Wang, Qiuchen, et autres
Publié: (2026)
par: Wang, Qiuchen, et autres
Publié: (2026)
VisRAG 2.0: Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation
par: Sun, Yubo, et autres
Publié: (2025)
par: Sun, Yubo, et autres
Publié: (2025)
The Devil is in the Prompts: Retrieval-Augmented Prompt Optimization for Text-to-Video Generation
par: Gao, Bingjie, et autres
Publié: (2025)
par: Gao, Bingjie, et autres
Publié: (2025)
VL-Cache: Sparsity and Modality-Aware KV Cache Compression for Vision-Language Model Inference Acceleration
par: Tu, Dezhan, et autres
Publié: (2024)
par: Tu, Dezhan, et autres
Publié: (2024)
Retrieval-Augmented Perception: High-Resolution Image Perception Meets Visual RAG
par: Wang, Wenbin, et autres
Publié: (2025)
par: Wang, Wenbin, et autres
Publié: (2025)
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization
par: Tao, Wei, et autres
Publié: (2026)
par: Tao, Wei, et autres
Publié: (2026)
mRAG: Elucidating the Design Space of Multi-modal Retrieval-Augmented Generation
par: Hu, Chan-Wei, et autres
Publié: (2025)
par: Hu, Chan-Wei, et autres
Publié: (2025)
Visual Late Chunking: An Empirical Study of Contextual Chunking for Efficient Visual Document Retrieval
par: Yan, Yibo, et autres
Publié: (2026)
par: Yan, Yibo, et autres
Publié: (2026)
LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference
par: Wan, Zhongwei, et autres
Publié: (2024)
par: Wan, Zhongwei, et autres
Publié: (2024)
Understanding Retrieval Robustness for Retrieval-Augmented Image Captioning
par: Li, Wenyan, et autres
Publié: (2024)
par: Li, Wenyan, et autres
Publié: (2024)
StreamKV: Streaming Video Question-Answering with Segment-based KV Cache Retrieval and Compression
par: Chen, Yilong, et autres
Publié: (2025)
par: Chen, Yilong, et autres
Publié: (2025)
NavRAG: Generating User Demand Instructions for Embodied Navigation through Retrieval-Augmented LLM
par: Wang, Zihan, et autres
Publié: (2025)
par: Wang, Zihan, et autres
Publié: (2025)
VDocRAG: Retrieval-Augmented Generation over Visually-Rich Documents
par: Tanaka, Ryota, et autres
Publié: (2025)
par: Tanaka, Ryota, et autres
Publié: (2025)
MegaRAG: Multimodal Knowledge Graph-Based Retrieval Augmented Generation
par: Hsiao, Chi-Hsiang, et autres
Publié: (2025)
par: Hsiao, Chi-Hsiang, et autres
Publié: (2025)
HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding
par: Zhang, Haowei, et autres
Publié: (2026)
par: Zhang, Haowei, et autres
Publié: (2026)
Text-Only Training for Image Captioning with Retrieval Augmentation and Modality Gap Correction
par: Fonseca, Rui, et autres
Publié: (2025)
par: Fonseca, Rui, et autres
Publié: (2025)
When RAG Hurts: Diagnosing and Mitigating Attention Distraction in Retrieval-Augmented LVLMs
par: Zhao, Beidi, et autres
Publié: (2026)
par: Zhao, Beidi, et autres
Publié: (2026)
LagKV: Lag-Relative Information of the KV Cache Tells Which Tokens Are Important
par: Liang, Manlai, et autres
Publié: (2025)
par: Liang, Manlai, et autres
Publié: (2025)
KV-Efficient VLA: A Method to Speed up Vision Language Models with RNN-Gated Chunked KV Cache
par: Xu, Wanshun, et autres
Publié: (2025)
par: Xu, Wanshun, et autres
Publié: (2025)
FilterRAG: Zero-Shot Informed Retrieval-Augmented Generation to Mitigate Hallucinations in VQA
par: Sarwar, Nobin
Publié: (2025)
par: Sarwar, Nobin
Publié: (2025)
VideoRAG: Retrieval-Augmented Generation over Video Corpus
par: Jeong, Soyeong, et autres
Publié: (2025)
par: Jeong, Soyeong, et autres
Publié: (2025)
PackCache: A Training-Free Acceleration Method for Unified Autoregressive Video Generation via Compact KV-Cache
par: Li, Kunyang, et autres
Publié: (2026)
par: Li, Kunyang, et autres
Publié: (2026)
Benchmarking Retrieval-Augmented Multimodal Generation for Document Question Answering
par: Dong, Kuicai, et autres
Publié: (2025)
par: Dong, Kuicai, et autres
Publié: (2025)
AIP: Subverting Retrieval-Augmented Generation via Adversarial Instructional Prompt
par: Chaturvedi, Saket S., et autres
Publié: (2025)
par: Chaturvedi, Saket S., et autres
Publié: (2025)
PartRAG: Retrieval-Augmented Part-Level 3D Generation and Editing
par: Li, Peize, et autres
Publié: (2026)
par: Li, Peize, et autres
Publié: (2026)
LayoutRAG: Retrieval-Augmented Model for Content-agnostic Conditional Layout Generation
par: Wu, Yuxuan, et autres
Publié: (2025)
par: Wu, Yuxuan, et autres
Publié: (2025)
Streaming Video Question-Answering with In-context Video KV-Cache Retrieval
par: Di, Shangzhe, et autres
Publié: (2025)
par: Di, Shangzhe, et autres
Publié: (2025)
AgenticOCR: Parsing Only What You Need for Efficient Retrieval-Augmented Generation
par: Wang, Zhengren, et autres
Publié: (2026)
par: Wang, Zhengren, et autres
Publié: (2026)
Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document Understanding
par: Gao, Sensen, et autres
Publié: (2025)
par: Gao, Sensen, et autres
Publié: (2025)
Text-Video Retrieval via Variational Multi-Modal Hypergraph Networks
par: Li, Qian, et autres
Publié: (2024)
par: Li, Qian, et autres
Publié: (2024)
ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoning Agents
par: Wang, Qiuchen, et autres
Publié: (2025)
par: Wang, Qiuchen, et autres
Publié: (2025)
AR-RAG: Autoregressive Retrieval Augmentation for Image Generation
par: Qi, Jingyuan, et autres
Publié: (2025)
par: Qi, Jingyuan, et autres
Publié: (2025)
Open Vocabulary Panoptic Segmentation With Retrieval Augmentation
par: Sadeq, Nafis, et autres
Publié: (2026)
par: Sadeq, Nafis, et autres
Publié: (2026)
Documents similaires
-
URPO: A Unified Reward & Policy Optimization Framework for Large Language Models
par: Lu, Songshuo, et autres
Publié: (2025) -
Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries
par: Wu, Yin, et autres
Publié: (2025) -
MobileRAG: Enhancing Mobile Agent with Retrieval-Augmented Generation
par: Loo, Gowen, et autres
Publié: (2025) -
LaB-RAG: Label Boosted Retrieval Augmented Generation for Radiology Report Generation
par: Song, Steven, et autres
Publié: (2024) -
Cross-modal RAG: Sub-dimensional Text-to-Image Retrieval-Augmented Generation
par: Zhu, Mengdan, et autres
Publié: (2025)