CacheFocus: Dynamic Cache Re-Positioning for Efficient Retrieval-Augmented Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Lee, Kun-Hui, Park, Eunhwan, Han, Donghoon, Na, Seung-Hoon |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MERLIN: Multimodal Embedding Refinement via LLM-based Iterative Navigation for Text-Video Retrieval-Rerank Pipeline
di: Han, Donghoon, et al.
Pubblicazione: (2024)
di: Han, Donghoon, et al.
Pubblicazione: (2024)
From Prefix Cache to Fusion RAG Cache: Accelerating LLM Inference in Retrieval-Augmented Generation
di: Wang, Jiahao, et al.
Pubblicazione: (2026)
di: Wang, Jiahao, et al.
Pubblicazione: (2026)
Cache-Craft: Managing Chunk-Caches for Efficient Retrieval-Augmented Generation
di: Agarwal, Shubham, et al.
Pubblicazione: (2025)
di: Agarwal, Shubham, et al.
Pubblicazione: (2025)
Visual Words Meet BM25: Sparse Auto-Encoder Visual Word Scoring for Image Retrieval
di: Han, Donghoon, et al.
Pubblicazione: (2026)
di: Han, Donghoon, et al.
Pubblicazione: (2026)
HeteroCache: A Dynamic Retrieval Approach to Heterogeneous KV Cache Compression for Long-Context LLM Inference
di: Shi, Zhiyuan, et al.
Pubblicazione: (2026)
di: Shi, Zhiyuan, et al.
Pubblicazione: (2026)
Enhancing Robustness of Retrieval-Augmented Language Models with In-Context Learning
di: Park, Seong-Il, et al.
Pubblicazione: (2024)
di: Park, Seong-Il, et al.
Pubblicazione: (2024)
XC-Cache: Cross-Attending to Cached Context for Efficient LLM Inference
di: Monteiro, João, et al.
Pubblicazione: (2024)
di: Monteiro, João, et al.
Pubblicazione: (2024)
Prompt Cache: Modular Attention Reuse for Low-Latency Inference
di: Gim, In, et al.
Pubblicazione: (2023)
di: Gim, In, et al.
Pubblicazione: (2023)
Parallel Key-Value Cache Fusion for Position Invariant RAG
di: Oh, Philhoon, et al.
Pubblicazione: (2025)
di: Oh, Philhoon, et al.
Pubblicazione: (2025)
Scaling Test-Time Inference with Policy-Optimized, Dynamic Retrieval-Augmented Generation via KV Caching and Decoding
di: Srinivas, Sakhinana Sagar, et al.
Pubblicazione: (2025)
di: Srinivas, Sakhinana Sagar, et al.
Pubblicazione: (2025)
FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference
di: Liu, Guangda, et al.
Pubblicazione: (2025)
di: Liu, Guangda, et al.
Pubblicazione: (2025)
Enhancing Cache-Augmented Generation (CAG) with Adaptive Contextual Compression for Scalable Knowledge Integration
di: Agrawal, Rishabh, et al.
Pubblicazione: (2025)
di: Agrawal, Rishabh, et al.
Pubblicazione: (2025)
Grounded Cache Routing for Retrieval-Augmented Generation: When Is It Safe to Reuse an Answer?
di: Shah, Syed Huma
Pubblicazione: (2026)
di: Shah, Syed Huma
Pubblicazione: (2026)
Conflict-Aware Soft Prompting for Retrieval-Augmented Generation
di: Choi, Eunseong, et al.
Pubblicazione: (2025)
di: Choi, Eunseong, et al.
Pubblicazione: (2025)
Lossless KV Cache Compression to 2%
di: Yang, Zhen, et al.
Pubblicazione: (2024)
di: Yang, Zhen, et al.
Pubblicazione: (2024)
Generative Caching for Structurally Similar Prompts and Responses
di: Chakraborty, Sarthak, et al.
Pubblicazione: (2025)
di: Chakraborty, Sarthak, et al.
Pubblicazione: (2025)
FeRG-LLM : Feature Engineering by Reason Generation Large Language Models
di: Ko, Jeonghyun, et al.
Pubblicazione: (2025)
di: Ko, Jeonghyun, et al.
Pubblicazione: (2025)
ROSAQ: Rotation-based Saliency-Aware Weight Quantization for Efficiently Compressing Large Language Models
di: Yoon, Junho, et al.
Pubblicazione: (2025)
di: Yoon, Junho, et al.
Pubblicazione: (2025)
TokenSelect: Efficient Long-Context Inference and Length Extrapolation for LLMs via Dynamic Token-Level KV Cache Selection
di: Wu, Wei, et al.
Pubblicazione: (2024)
di: Wu, Wei, et al.
Pubblicazione: (2024)
Deliberation in Latent Space via Differentiable Cache Augmentation
di: Liu, Luyang, et al.
Pubblicazione: (2024)
di: Liu, Luyang, et al.
Pubblicazione: (2024)
PEAR: Position-Embedding-Agnostic Attention Re-weighting Enhances Retrieval-Augmented Generation with Zero Inference Overhead
di: Tan, Tao, et al.
Pubblicazione: (2024)
di: Tan, Tao, et al.
Pubblicazione: (2024)
LaCache: Ladder-Shaped KV Caching for Efficient Long-Context Modeling of Large Language Models
di: Shi, Dachuan, et al.
Pubblicazione: (2025)
di: Shi, Dachuan, et al.
Pubblicazione: (2025)
Lookahead Q-Cache: Achieving More Consistent KV Cache Eviction via Pseudo Query
di: Wang, Yixuan, et al.
Pubblicazione: (2025)
di: Wang, Yixuan, et al.
Pubblicazione: (2025)
EXIT: Context-Aware Extractive Compression for Enhancing Retrieval-Augmented Generation
di: Hwang, Taeho, et al.
Pubblicazione: (2024)
di: Hwang, Taeho, et al.
Pubblicazione: (2024)
Utility-Focused LLM Annotation for Retrieval and Retrieval-Augmented Generation
di: Zhang, Hengran, et al.
Pubblicazione: (2025)
di: Zhang, Hengran, et al.
Pubblicazione: (2025)
Q-Filters: Leveraging QK Geometry for Efficient KV Cache Compression
di: Godey, Nathan, et al.
Pubblicazione: (2025)
di: Godey, Nathan, et al.
Pubblicazione: (2025)
TableCache: Primary Foreign Key Guided KV Cache Precomputation for Low Latency Text-to-SQL
di: Su, Jinbo, et al.
Pubblicazione: (2026)
di: Su, Jinbo, et al.
Pubblicazione: (2026)
AgenticCache: Cache-Driven Asynchronous Planning for Embodied AI Agents
di: Kim, Hojoon, et al.
Pubblicazione: (2026)
di: Kim, Hojoon, et al.
Pubblicazione: (2026)
Cacheback: Speculative Decoding With Nothing But Cache
di: Ma, Zhiyao, et al.
Pubblicazione: (2025)
di: Ma, Zhiyao, et al.
Pubblicazione: (2025)
SUGAR: Leveraging Contextual Confidence for Smarter Retrieval
di: Zubkova, Hanna, et al.
Pubblicazione: (2025)
di: Zubkova, Hanna, et al.
Pubblicazione: (2025)
Beyond KV Caching: Shared Attention for Efficient LLMs
di: Liao, Bingli, et al.
Pubblicazione: (2024)
di: Liao, Bingli, et al.
Pubblicazione: (2024)
PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
di: Cai, Zefan, et al.
Pubblicazione: (2024)
di: Cai, Zefan, et al.
Pubblicazione: (2024)
dLLM-Cache: Accelerating Diffusion Large Language Models with Adaptive Caching
di: Liu, Zhiyuan, et al.
Pubblicazione: (2025)
di: Liu, Zhiyuan, et al.
Pubblicazione: (2025)
MiniCache: KV Cache Compression in Depth Dimension for Large Language Models
di: Liu, Akide, et al.
Pubblicazione: (2024)
di: Liu, Akide, et al.
Pubblicazione: (2024)
Not All Heads Matter: A Head-Level KV Cache Compression Method with Integrated Retrieval and Reasoning
di: Fu, Yu, et al.
Pubblicazione: (2024)
di: Fu, Yu, et al.
Pubblicazione: (2024)
Towards Threshold-Free KV Cache Pruning
di: Ni, Xuanfan, et al.
Pubblicazione: (2025)
di: Ni, Xuanfan, et al.
Pubblicazione: (2025)
KV Cache Steering for Controlling Frozen LLMs
di: Belitsky, Max, et al.
Pubblicazione: (2025)
di: Belitsky, Max, et al.
Pubblicazione: (2025)
ZSMerge: Zero-Shot KV Cache Compression for Memory-Efficient Long-Context LLMs
di: Liu, Xin, et al.
Pubblicazione: (2025)
di: Liu, Xin, et al.
Pubblicazione: (2025)
KVShare: An LLM Service System with Efficient and Effective Multi-Tenant KV Cache Reuse
di: Yang, Huan, et al.
Pubblicazione: (2025)
di: Yang, Huan, et al.
Pubblicazione: (2025)
OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference
di: Gu, Yuzhe, et al.
Pubblicazione: (2025)
di: Gu, Yuzhe, et al.
Pubblicazione: (2025)
Documenti analoghi
-
MERLIN: Multimodal Embedding Refinement via LLM-based Iterative Navigation for Text-Video Retrieval-Rerank Pipeline
di: Han, Donghoon, et al.
Pubblicazione: (2024) -
From Prefix Cache to Fusion RAG Cache: Accelerating LLM Inference in Retrieval-Augmented Generation
di: Wang, Jiahao, et al.
Pubblicazione: (2026) -
Cache-Craft: Managing Chunk-Caches for Efficient Retrieval-Augmented Generation
di: Agarwal, Shubham, et al.
Pubblicazione: (2025) -
Visual Words Meet BM25: Sparse Auto-Encoder Visual Word Scoring for Image Retrieval
di: Han, Donghoon, et al.
Pubblicazione: (2026) -
HeteroCache: A Dynamic Retrieval Approach to Heterogeneous KV Cache Compression for Long-Context LLM Inference
di: Shi, Zhiyuan, et al.
Pubblicazione: (2026)