EFIM: Efficient Serving of LLMs for Infilling Tasks with Improved KV Cache Reuse
Fuente:
arXiv
Guardado en:
| Autores principales: | Guo, Tianyu, Dong, Hande, Leng, Yichong, Liu, Feng, Lin, Cheater, Xiao, Nong, Zhang, Xianwei |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs
por: Zhou, Xiabin, et al.
Publicado: (2024)
por: Zhou, Xiabin, et al.
Publicado: (2024)
KVShare: An LLM Service System with Efficient and Effective Multi-Tenant KV Cache Reuse
por: Yang, Huan, et al.
Publicado: (2025)
por: Yang, Huan, et al.
Publicado: (2025)
KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse
por: Yang, Jingbo, et al.
Publicado: (2025)
por: Yang, Jingbo, et al.
Publicado: (2025)
PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
por: Cai, Zefan, et al.
Publicado: (2024)
por: Cai, Zefan, et al.
Publicado: (2024)
WindowKV: Task-Adaptive Group-Wise KV Cache Window Selection for Efficient LLM Inference
por: Zuo, Youhui, et al.
Publicado: (2025)
por: Zuo, Youhui, et al.
Publicado: (2025)
FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference
por: Liu, Guangda, et al.
Publicado: (2025)
por: Liu, Guangda, et al.
Publicado: (2025)
Beyond Homogeneous Attention: Memory-Efficient LLMs via Fourier-Approximated KV Cache
por: Liu, Xiaoran, et al.
Publicado: (2025)
por: Liu, Xiaoran, et al.
Publicado: (2025)
Task-KV: Task-aware KV Cache Optimization via Semantic Differentiation of Attention Heads
por: He, Xingyang, et al.
Publicado: (2025)
por: He, Xingyang, et al.
Publicado: (2025)
CacheClip: Accelerating RAG with Effective KV Cache Reuse
por: Yang, Bin, et al.
Publicado: (2025)
por: Yang, Bin, et al.
Publicado: (2025)
Beyond KV Caching: Shared Attention for Efficient LLMs
por: Liao, Bingli, et al.
Publicado: (2024)
por: Liao, Bingli, et al.
Publicado: (2024)
Taming the Fragility of KV Cache Eviction in LLM Inference
por: Feng, Yuan, et al.
Publicado: (2025)
por: Feng, Yuan, et al.
Publicado: (2025)
ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference
por: Liu, Xiang, et al.
Publicado: (2025)
por: Liu, Xiang, et al.
Publicado: (2025)
In-context KV-Cache Eviction for LLMs via Attention-Gate
por: Zeng, Zihao, et al.
Publicado: (2024)
por: Zeng, Zihao, et al.
Publicado: (2024)
Adaptive KV Cache Reuse for Fast Long-Context LLM Serving
por: li, Fei, et al.
Publicado: (2026)
por: li, Fei, et al.
Publicado: (2026)
ZSMerge: Zero-Shot KV Cache Compression for Memory-Efficient Long-Context LLMs
por: Liu, Xin, et al.
Publicado: (2025)
por: Liu, Xin, et al.
Publicado: (2025)
Model Tells You Where to Merge: Adaptive KV Cache Merging for LLMs on Long-Context Tasks
por: Wang, Zheng, et al.
Publicado: (2024)
por: Wang, Zheng, et al.
Publicado: (2024)
EchoKV: Efficient KV Cache Compression via Similarity-Based Reconstruction
por: Ji, Shiyu, et al.
Publicado: (2026)
por: Ji, Shiyu, et al.
Publicado: (2026)
CriticalKV: Optimizing KV Cache Eviction from an Output Perturbation Perspective
por: Feng, Yuan, et al.
Publicado: (2025)
por: Feng, Yuan, et al.
Publicado: (2025)
Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs
por: Ge, Suyu, et al.
Publicado: (2023)
por: Ge, Suyu, et al.
Publicado: (2023)
KV Cache Steering for Controlling Frozen LLMs
por: Belitsky, Max, et al.
Publicado: (2025)
por: Belitsky, Max, et al.
Publicado: (2025)
DroidSpeak: KV Cache Sharing for Cross-LLM Communication and Multi-LLM Serving
por: Liu, Yuhan, et al.
Publicado: (2024)
por: Liu, Yuhan, et al.
Publicado: (2024)
KV-CoRE: Benchmarking Data-Dependent Low-Rank Compressibility of KV-Caches in LLMs
por: Chen, Jian, et al.
Publicado: (2026)
por: Chen, Jian, et al.
Publicado: (2026)
TreeKV: Smooth Key-Value Cache Compression with Tree Structures
por: He, Ziwei, et al.
Publicado: (2025)
por: He, Ziwei, et al.
Publicado: (2025)
HyperRAG: Enhancing Quality-Efficiency Tradeoffs in Retrieval-Augmented Generation with Reranker KV-Cache Reuse
por: An, Yuwei, et al.
Publicado: (2025)
por: An, Yuwei, et al.
Publicado: (2025)
SemantiCache: Efficient KV Cache Compression via Semantic Chunking and Clustered Merging
por: Wu, Shunlong, et al.
Publicado: (2026)
por: Wu, Shunlong, et al.
Publicado: (2026)
KV Cache Offloading for Context-Intensive Tasks
por: Bocharnikov, Andrey, et al.
Publicado: (2026)
por: Bocharnikov, Andrey, et al.
Publicado: (2026)
Crystal-KV: Efficient KV Cache Management for Chain-of-Thought LLMs via Answer-First Principle
por: Wang, Zihan, et al.
Publicado: (2026)
por: Wang, Zihan, et al.
Publicado: (2026)
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
por: Feng, Yuan, et al.
Publicado: (2024)
por: Feng, Yuan, et al.
Publicado: (2024)
gLLM: Global Balanced Pipeline Parallelism System for Distributed LLM Serving with Token Throttling
por: Guo, Tianyu, et al.
Publicado: (2025)
por: Guo, Tianyu, et al.
Publicado: (2025)
AhaKV: Adaptive Holistic Attention-Driven KV Cache Eviction for Efficient Inference of Large Language Models
por: Gu, Yifeng, et al.
Publicado: (2025)
por: Gu, Yifeng, et al.
Publicado: (2025)
R-KV: Redundancy-aware KV Cache Compression for Reasoning Models
por: Cai, Zefan, et al.
Publicado: (2025)
por: Cai, Zefan, et al.
Publicado: (2025)
PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs
por: Liu, Tengxuan, et al.
Publicado: (2025)
por: Liu, Tengxuan, et al.
Publicado: (2025)
RazorAttention: Efficient KV Cache Compression Through Retrieval Heads
por: Tang, Hanlin, et al.
Publicado: (2024)
por: Tang, Hanlin, et al.
Publicado: (2024)
NestedKV: Nested Memory Routing for Long-Context KV Cache Compression
por: Chen, Hong, et al.
Publicado: (2026)
por: Chen, Hong, et al.
Publicado: (2026)
G-KV: Decoding-Time KV Cache Eviction with Global Attention
por: Liao, Mengqi, et al.
Publicado: (2025)
por: Liao, Mengqi, et al.
Publicado: (2025)
GUI-KV: Efficient GUI Agents via KV Cache with Spatio-Temporal Awareness
por: Huang, Kung-Hsiang, et al.
Publicado: (2025)
por: Huang, Kung-Hsiang, et al.
Publicado: (2025)
PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference
por: Yang, Dongjie, et al.
Publicado: (2024)
por: Yang, Dongjie, et al.
Publicado: (2024)
LaCache: Ladder-Shaped KV Caching for Efficient Long-Context Modeling of Large Language Models
por: Shi, Dachuan, et al.
Publicado: (2025)
por: Shi, Dachuan, et al.
Publicado: (2025)
MEDA: Dynamic KV Cache Allocation for Efficient Multimodal Long-Context Inference
por: Wan, Zhongwei, et al.
Publicado: (2025)
por: Wan, Zhongwei, et al.
Publicado: (2025)
$A^3$: Attention-Aware Accurate KV Cache Fusion for Fast Large Language Model Serving
por: Zhou, Yuechi, et al.
Publicado: (2025)
por: Zhou, Yuechi, et al.
Publicado: (2025)
Ejemplares similares
-
DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs
por: Zhou, Xiabin, et al.
Publicado: (2024) -
KVShare: An LLM Service System with Efficient and Effective Multi-Tenant KV Cache Reuse
por: Yang, Huan, et al.
Publicado: (2025) -
KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse
por: Yang, Jingbo, et al.
Publicado: (2025) -
PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
por: Cai, Zefan, et al.
Publicado: (2024) -
WindowKV: Task-Adaptive Group-Wise KV Cache Window Selection for Efficient LLM Inference
por: Zuo, Youhui, et al.
Publicado: (2025)