NACL: A General and Effective KV Cache Eviction Framework for LLMs at Inference Time
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Yilong, Wang, Guoxia, Shang, Junyuan, Cui, Shiyao, Zhang, Zhenyu, Liu, Tingwen, Wang, Shuohuan, Sun, Yu, Yu, Dianhai, Wu, Hua |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DHA: Learning Decoupled-Head Attention from Transformer Checkpoints via Adaptive Heads Fusion
von: Chen, Yilong, et al.
Veröffentlicht: (2024)
von: Chen, Yilong, et al.
Veröffentlicht: (2024)
Inner Thinking Transformer: Leveraging Dynamic Depth Scaling to Foster Adaptive Internal Thinking
von: Chen, Yilong, et al.
Veröffentlicht: (2025)
von: Chen, Yilong, et al.
Veröffentlicht: (2025)
Mixture of Hidden-Dimensions Transformer
von: Chen, Yilong, et al.
Veröffentlicht: (2024)
von: Chen, Yilong, et al.
Veröffentlicht: (2024)
Mixture of Universal Experts: Scaling Virtual Width via Depth-Width Transformation
von: Chen, Yilong, et al.
Veröffentlicht: (2026)
von: Chen, Yilong, et al.
Veröffentlicht: (2026)
Sparse Growing Transformer: Training-Time Sparse Depth Allocation via Progressive Attention Looping
von: Chen, Yao, et al.
Veröffentlicht: (2026)
von: Chen, Yao, et al.
Veröffentlicht: (2026)
MomentKV: Closing the Directional Gap in KV Cache Eviction for Long-Context Inference
von: Li, Yu, et al.
Veröffentlicht: (2026)
von: Li, Yu, et al.
Veröffentlicht: (2026)
Taming the Fragility of KV Cache Eviction in LLM Inference
von: Feng, Yuan, et al.
Veröffentlicht: (2025)
von: Feng, Yuan, et al.
Veröffentlicht: (2025)
G-KV: Decoding-Time KV Cache Eviction with Global Attention
von: Liao, Mengqi, et al.
Veröffentlicht: (2025)
von: Liao, Mengqi, et al.
Veröffentlicht: (2025)
MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context Inference
von: Li, Kunxi, et al.
Veröffentlicht: (2025)
von: Li, Kunxi, et al.
Veröffentlicht: (2025)
VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator
von: Wang, Zhican, et al.
Veröffentlicht: (2025)
von: Wang, Zhican, et al.
Veröffentlicht: (2025)
VideoAR: Autoregressive Video Generation via Next-Frame & Scale Prediction
von: Ji, Longbin, et al.
Veröffentlicht: (2026)
von: Ji, Longbin, et al.
Veröffentlicht: (2026)
FFT: Towards Harmlessness Evaluation and Analysis for LLMs with Factuality, Fairness, Toxicity
von: Cui, Shiyao, et al.
Veröffentlicht: (2023)
von: Cui, Shiyao, et al.
Veröffentlicht: (2023)
MPCache: MPC-Friendly KV Cache Eviction for Efficient Private LLM Inference
von: Zeng, Wenxuan, et al.
Veröffentlicht: (2025)
von: Zeng, Wenxuan, et al.
Veröffentlicht: (2025)
LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction for Efficient Long-Context Inference
von: Wang, Guangtao, et al.
Veröffentlicht: (2025)
von: Wang, Guangtao, et al.
Veröffentlicht: (2025)
RetentiveKV: State-Space Memory for Uncertainty-Aware Multimodal KV Cache Eviction
von: Liu, Sihao, et al.
Veröffentlicht: (2026)
von: Liu, Sihao, et al.
Veröffentlicht: (2026)
In-context KV-Cache Eviction for LLMs via Attention-Gate
von: Zeng, Zihao, et al.
Veröffentlicht: (2024)
von: Zeng, Zihao, et al.
Veröffentlicht: (2024)
Reformulating KV Cache Eviction Problem for Long-Context LLM Inference
von: Mai, Tho, et al.
Veröffentlicht: (2026)
von: Mai, Tho, et al.
Veröffentlicht: (2026)
ReST-KV: Robust KV Cache Eviction with Layer-wise Output Reconstruction and Spatial-Temporal Smoothing
von: An, Yongqi, et al.
Veröffentlicht: (2026)
von: An, Yongqi, et al.
Veröffentlicht: (2026)
Elastic MoE: Unlocking the Inference-Time Scalability of Mixture-of-Experts
von: Gu, Naibin, et al.
Veröffentlicht: (2025)
von: Gu, Naibin, et al.
Veröffentlicht: (2025)
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
von: Feng, Yuan, et al.
Veröffentlicht: (2024)
von: Feng, Yuan, et al.
Veröffentlicht: (2024)
AudioKV: KV Cache Eviction in Efficient Large Audio Language Models
von: Wang, Yuxuan, et al.
Veröffentlicht: (2026)
von: Wang, Yuxuan, et al.
Veröffentlicht: (2026)
LookaheadKV: Fast and Accurate KV Cache Eviction by Glimpsing into the Future without Generation
von: Ahn, Jinwoo, et al.
Veröffentlicht: (2026)
von: Ahn, Jinwoo, et al.
Veröffentlicht: (2026)
Beyond Token Eviction: Mixed-Dimension Budget Allocation for Efficient KV Cache Compression
von: Miao, Ruijie, et al.
Veröffentlicht: (2026)
von: Miao, Ruijie, et al.
Veröffentlicht: (2026)
AhaKV: Adaptive Holistic Attention-Driven KV Cache Eviction for Efficient Inference of Large Language Models
von: Gu, Yifeng, et al.
Veröffentlicht: (2025)
von: Gu, Yifeng, et al.
Veröffentlicht: (2025)
LAVa: Layer-wise KV Cache Eviction with Dynamic Budget Allocation
von: Shen, Yiqun, et al.
Veröffentlicht: (2025)
von: Shen, Yiqun, et al.
Veröffentlicht: (2025)
V-ITI: Mitigating Hallucinations in Multimodal Large Language Models via Visual Inference-Time Intervention
von: Sun, Nan, et al.
Veröffentlicht: (2025)
von: Sun, Nan, et al.
Veröffentlicht: (2025)
CriticalKV: Optimizing KV Cache Eviction from an Output Perturbation Perspective
von: Feng, Yuan, et al.
Veröffentlicht: (2025)
von: Feng, Yuan, et al.
Veröffentlicht: (2025)
ForesightKV: Optimizing KV Cache Eviction for Reasoning Models by Learning Long-Term Contribution
von: Dong, Zican, et al.
Veröffentlicht: (2026)
von: Dong, Zican, et al.
Veröffentlicht: (2026)
PagedEviction: Structured Block-wise KV Cache Pruning for Efficient Large Language Model Inference
von: Chitty-Venkata, Krishna Teja, et al.
Veröffentlicht: (2025)
von: Chitty-Venkata, Krishna Teja, et al.
Veröffentlicht: (2025)
Lookahead Q-Cache: Achieving More Consistent KV Cache Eviction via Pseudo Query
von: Wang, Yixuan, et al.
Veröffentlicht: (2025)
von: Wang, Yixuan, et al.
Veröffentlicht: (2025)
CAKE: Cascading and Adaptive KV Cache Eviction with Layer Preferences
von: Qin, Ziran, et al.
Veröffentlicht: (2025)
von: Qin, Ziran, et al.
Veröffentlicht: (2025)
IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM Inference
von: Yang, Xintong, et al.
Veröffentlicht: (2026)
von: Yang, Xintong, et al.
Veröffentlicht: (2026)
SABlock: Semantic-Aware KV Cache Eviction with Adaptive Compression Block Size
von: Chen, Jinhan, et al.
Veröffentlicht: (2025)
von: Chen, Jinhan, et al.
Veröffentlicht: (2025)
GraphKV: Breaking the Static Selection Paradigm with Graph-Based KV Cache Eviction
von: Li, Xuelin, et al.
Veröffentlicht: (2025)
von: Li, Xuelin, et al.
Veröffentlicht: (2025)
CAOTE: KV Cache Selection for LLMs via Attention Output Error-Based Token Eviction
von: Goel, Raghavv, et al.
Veröffentlicht: (2025)
von: Goel, Raghavv, et al.
Veröffentlicht: (2025)
Judge Q: Trainable Queries for Optimized Information Retention in KV Cache Eviction
von: Liu, Yijun, et al.
Veröffentlicht: (2025)
von: Liu, Yijun, et al.
Veröffentlicht: (2025)
MoR: Mixture of Ranks for Low-Rank Adaptation Tuning
von: Tang, Chuanyu, et al.
Veröffentlicht: (2024)
von: Tang, Chuanyu, et al.
Veröffentlicht: (2024)
EvolKV: Evolutionary KV Cache Compression for LLM Inference
von: Yu, Bohan, et al.
Veröffentlicht: (2025)
von: Yu, Bohan, et al.
Veröffentlicht: (2025)
Hierarchical Adaptive Eviction for KV Cache Management in Multimodal Language Models
von: Ma, Xindian, et al.
Veröffentlicht: (2026)
von: Ma, Xindian, et al.
Veröffentlicht: (2026)
A Simple Plug-in for Improving Eviction-Based KV Cache Compression
von: Lin, Yuping, et al.
Veröffentlicht: (2026)
von: Lin, Yuping, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
DHA: Learning Decoupled-Head Attention from Transformer Checkpoints via Adaptive Heads Fusion
von: Chen, Yilong, et al.
Veröffentlicht: (2024) -
Inner Thinking Transformer: Leveraging Dynamic Depth Scaling to Foster Adaptive Internal Thinking
von: Chen, Yilong, et al.
Veröffentlicht: (2025) -
Mixture of Hidden-Dimensions Transformer
von: Chen, Yilong, et al.
Veröffentlicht: (2024) -
Mixture of Universal Experts: Scaling Virtual Width via Depth-Width Transformation
von: Chen, Yilong, et al.
Veröffentlicht: (2026) -
Sparse Growing Transformer: Training-Time Sparse Depth Allocation via Progressive Attention Looping
von: Chen, Yao, et al.
Veröffentlicht: (2026)