SpeCache: Speculative Key-Value Caching for Efficient Generation of LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Jie, Shibo, Tang, Yehui, Han, Kai, Deng, Zhi-Hong, Han, Jing |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Memory-Space Visual Prompting for Efficient Vision-Language Fine-Tuning
por: Jie, Shibo, et al.
Publicado: (2024)
por: Jie, Shibo, et al.
Publicado: (2024)
Mixture of Lookup Experts
por: Jie, Shibo, et al.
Publicado: (2025)
por: Jie, Shibo, et al.
Publicado: (2025)
EMS-SD: Efficient Multi-sample Speculative Decoding for Accelerating Large Language Models
por: Ni, Yunsheng, et al.
Publicado: (2024)
por: Ni, Yunsheng, et al.
Publicado: (2024)
SpeCa: Accelerating Diffusion Transformers with Speculative Feature Caching
por: Liu, Jiacheng, et al.
Publicado: (2025)
por: Liu, Jiacheng, et al.
Publicado: (2025)
Understanding the Physics of Key-Value Cache Compression for LLMs through Attention Dynamics
por: Ananthanarayanan, Samhruth, et al.
Publicado: (2026)
por: Ananthanarayanan, Samhruth, et al.
Publicado: (2026)
CacheFocus: Dynamic Cache Re-Positioning for Efficient Retrieval-Augmented Generation
por: Lee, Kun-Hui, et al.
Publicado: (2025)
por: Lee, Kun-Hui, et al.
Publicado: (2025)
SCOPE: Optimizing Key-Value Cache Compression in Long-context Generation
por: Wu, Jialong, et al.
Publicado: (2024)
por: Wu, Jialong, et al.
Publicado: (2024)
FlowKV: Enhancing Multi-Turn Conversational Coherence in LLMs via Isolated Key-Value Cache Management
por: Liu, Xiang, et al.
Publicado: (2025)
por: Liu, Xiang, et al.
Publicado: (2025)
RazorAttention: Efficient KV Cache Compression Through Retrieval Heads
por: Tang, Hanlin, et al.
Publicado: (2024)
por: Tang, Hanlin, et al.
Publicado: (2024)
Learning to Evict from Key-Value Cache
por: Moschella, Luca, et al.
Publicado: (2026)
por: Moschella, Luca, et al.
Publicado: (2026)
SONIC: Segmented Optimized Nexus for Information Compression in Key-Value Caching
por: Chen, Hong, et al.
Publicado: (2026)
por: Chen, Hong, et al.
Publicado: (2026)
Cacheback: Speculative Decoding With Nothing But Cache
por: Ma, Zhiyao, et al.
Publicado: (2025)
por: Ma, Zhiyao, et al.
Publicado: (2025)
TableCache: Primary Foreign Key Guided KV Cache Precomputation for Low Latency Text-to-SQL
por: Su, Jinbo, et al.
Publicado: (2026)
por: Su, Jinbo, et al.
Publicado: (2026)
Kangaroo: Lossless Self-Speculative Decoding via Double Early Exiting
por: Liu, Fangcheng, et al.
Publicado: (2024)
por: Liu, Fangcheng, et al.
Publicado: (2024)
TreeKV: Smooth Key-Value Cache Compression with Tree Structures
por: He, Ziwei, et al.
Publicado: (2025)
por: He, Ziwei, et al.
Publicado: (2025)
Parallel Key-Value Cache Fusion for Position Invariant RAG
por: Oh, Philhoon, et al.
Publicado: (2025)
por: Oh, Philhoon, et al.
Publicado: (2025)
Homogeneous Keys, Heterogeneous Values: Exploiting Local KV Cache Asymmetry for Long-Context LLMs
por: Cui, Wanyun, et al.
Publicado: (2025)
por: Cui, Wanyun, et al.
Publicado: (2025)
Key, Value, Compress: A Systematic Exploration of KV Cache Compression Techniques
por: Javidnia, Neusha, et al.
Publicado: (2025)
por: Javidnia, Neusha, et al.
Publicado: (2025)
CSR:Achieving 1 Bit Key-Value Cache via Sparse Representation
por: Zhang, Hongxuan, et al.
Publicado: (2024)
por: Zhang, Hongxuan, et al.
Publicado: (2024)
Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
por: Brandon, William, et al.
Publicado: (2024)
por: Brandon, William, et al.
Publicado: (2024)
When Do LLMs Reason? A Dynamical Systems View via Entropy Phase Transitions
por: Xia, Wei, et al.
Publicado: (2026)
por: Xia, Wei, et al.
Publicado: (2026)
SemantiCache: Efficient KV Cache Compression via Semantic Chunking and Clustered Merging
por: Wu, Shunlong, et al.
Publicado: (2026)
por: Wu, Shunlong, et al.
Publicado: (2026)
KV-Runahead: Scalable Causal LLM Inference by Parallel Key-Value Cache Generation
por: Cho, Minsik, et al.
Publicado: (2024)
por: Cho, Minsik, et al.
Publicado: (2024)
From Prefix Cache to Fusion RAG Cache: Accelerating LLM Inference in Retrieval-Augmented Generation
por: Wang, Jiahao, et al.
Publicado: (2026)
por: Wang, Jiahao, et al.
Publicado: (2026)
PocketLLM: Ultimate Compression of Large Language Models via Meta Networks
por: Tian, Ye, et al.
Publicado: (2025)
por: Tian, Ye, et al.
Publicado: (2025)
SKVQ: Sliding-window Key and Value Cache Quantization for Large Language Models
por: Duanmu, Haojie, et al.
Publicado: (2024)
por: Duanmu, Haojie, et al.
Publicado: (2024)
EpiCache: Episodic KV Cache Management for Long-Term Conversation on Resource-Constrained Environments
por: Kim, Minsoo, et al.
Publicado: (2025)
por: Kim, Minsoo, et al.
Publicado: (2025)
d$^2$Cache: Accelerating Diffusion-Based LLMs via Dual Adaptive Caching
por: Jiang, Yuchu, et al.
Publicado: (2025)
por: Jiang, Yuchu, et al.
Publicado: (2025)
InstCache: A Predictive Cache for LLM Serving
por: Zou, Longwei, et al.
Publicado: (2024)
por: Zou, Longwei, et al.
Publicado: (2024)
Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs
por: Ge, Suyu, et al.
Publicado: (2023)
por: Ge, Suyu, et al.
Publicado: (2023)
Token Compensator: Altering Inference Cost of Vision Transformer without Re-Tuning
por: Jie, Shibo, et al.
Publicado: (2024)
por: Jie, Shibo, et al.
Publicado: (2024)
WeightedKV: Attention Scores Weighted Key-Value Cache Merging for Large Language Models
por: Yuan, Jian, et al.
Publicado: (2025)
por: Yuan, Jian, et al.
Publicado: (2025)
ZSMerge: Zero-Shot KV Cache Compression for Memory-Efficient Long-Context LLMs
por: Liu, Xin, et al.
Publicado: (2025)
por: Liu, Xin, et al.
Publicado: (2025)
SQuat: Subspace-orthogonal KV Cache Quantization
por: Wang, Hao, et al.
Publicado: (2025)
por: Wang, Hao, et al.
Publicado: (2025)
Cache-Craft: Managing Chunk-Caches for Efficient Retrieval-Augmented Generation
por: Agarwal, Shubham, et al.
Publicado: (2025)
por: Agarwal, Shubham, et al.
Publicado: (2025)
ThinK: Thinner Key Cache by Query-Driven Pruning
por: Xu, Yuhui, et al.
Publicado: (2024)
por: Xu, Yuhui, et al.
Publicado: (2024)
Beyond KV Caching: Shared Attention for Efficient LLMs
por: Liao, Bingli, et al.
Publicado: (2024)
por: Liao, Bingli, et al.
Publicado: (2024)
In-context KV-Cache Eviction for LLMs via Attention-Gate
por: Zeng, Zihao, et al.
Publicado: (2024)
por: Zeng, Zihao, et al.
Publicado: (2024)
XC-Cache: Cross-Attending to Cached Context for Efficient LLM Inference
por: Monteiro, João, et al.
Publicado: (2024)
por: Monteiro, João, et al.
Publicado: (2024)
GRKV: Global Regression for Training-Free KV Cache Compression in Long-Context LLMs
por: Peng, Junjie, et al.
Publicado: (2026)
por: Peng, Junjie, et al.
Publicado: (2026)
Ejemplares similares
-
Memory-Space Visual Prompting for Efficient Vision-Language Fine-Tuning
por: Jie, Shibo, et al.
Publicado: (2024) -
Mixture of Lookup Experts
por: Jie, Shibo, et al.
Publicado: (2025) -
EMS-SD: Efficient Multi-sample Speculative Decoding for Accelerating Large Language Models
por: Ni, Yunsheng, et al.
Publicado: (2024) -
SpeCa: Accelerating Diffusion Transformers with Speculative Feature Caching
por: Liu, Jiacheng, et al.
Publicado: (2025) -
Understanding the Physics of Key-Value Cache Compression for LLMs through Attention Dynamics
por: Ananthanarayanan, Samhruth, et al.
Publicado: (2026)