Towards Threshold-Free KV Cache Pruning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ni, Xuanfan, Xu, Liyan, Lyu, Chenyang, Wang, Longyue, Yu, Mo, Liu, Lemao, Meng, Fandong, Zhou, Jie, Li, Piji |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Systematic Evaluation of Large Language Models for Natural Language Generation Tasks
von: Ni, Xuanfan, et al.
Veröffentlicht: (2024)
von: Ni, Xuanfan, et al.
Veröffentlicht: (2024)
Marco-o1 v2: Towards Widening The Distillation Bottleneck for Reasoning Models
von: Yin, Huifeng, et al.
Veröffentlicht: (2025)
von: Yin, Huifeng, et al.
Veröffentlicht: (2025)
DepthKV: Layer-Dependent KV Cache Pruning for Long-Context LLM Inference
von: Dehghanighobadi, Zahra, et al.
Veröffentlicht: (2026)
von: Dehghanighobadi, Zahra, et al.
Veröffentlicht: (2026)
KVzap: Fast, Adaptive, and Faithful KV Cache Pruning
von: Jegou, Simon, et al.
Veröffentlicht: (2026)
von: Jegou, Simon, et al.
Veröffentlicht: (2026)
FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference
von: Liu, Guangda, et al.
Veröffentlicht: (2025)
von: Liu, Guangda, et al.
Veröffentlicht: (2025)
SparK: Query-Aware Unstructured Sparsity with Recoverable KV Cache Channel Pruning
von: Liao, Huanxuan, et al.
Veröffentlicht: (2025)
von: Liao, Huanxuan, et al.
Veröffentlicht: (2025)
CSCD-NS: a Chinese Spelling Check Dataset for Native Speakers
von: Hu, Yong, et al.
Veröffentlicht: (2022)
von: Hu, Yong, et al.
Veröffentlicht: (2022)
DeepTrans: Deep Reasoning Translation via Reinforcement Learning
von: Wang, Jiaan, et al.
Veröffentlicht: (2025)
von: Wang, Jiaan, et al.
Veröffentlicht: (2025)
ExTrans: Multilingual Deep Reasoning Translation via Exemplar-Enhanced Reinforcement Learning
von: Wang, Jiaan, et al.
Veröffentlicht: (2025)
von: Wang, Jiaan, et al.
Veröffentlicht: (2025)
OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference
von: Gu, Yuzhe, et al.
Veröffentlicht: (2025)
von: Gu, Yuzhe, et al.
Veröffentlicht: (2025)
EAG: Extract and Generate Multi-way Aligned Corpus for Complete Multi-lingual Neural Machine Translation
von: Xu, Yulin, et al.
Veröffentlicht: (2022)
von: Xu, Yulin, et al.
Veröffentlicht: (2022)
On the token distance modeling ability of higher RoPE attention dimension
von: Hong, Xiangyu, et al.
Veröffentlicht: (2024)
von: Hong, Xiangyu, et al.
Veröffentlicht: (2024)
Lookahead Q-Cache: Achieving More Consistent KV Cache Eviction via Pseudo Query
von: Wang, Yixuan, et al.
Veröffentlicht: (2025)
von: Wang, Yixuan, et al.
Veröffentlicht: (2025)
R-KV: Redundancy-aware KV Cache Compression for Reasoning Models
von: Cai, Zefan, et al.
Veröffentlicht: (2025)
von: Cai, Zefan, et al.
Veröffentlicht: (2025)
Beyond Next Token Prediction: Patch-Level Training for Large Language Models
von: Shao, Chenze, et al.
Veröffentlicht: (2024)
von: Shao, Chenze, et al.
Veröffentlicht: (2024)
ManifoldKV: Training-Free KV Cache Compression via Euclidean Outlier Detection
von: Datta, Debajyoti, et al.
Veröffentlicht: (2026)
von: Datta, Debajyoti, et al.
Veröffentlicht: (2026)
ThinK: Thinner Key Cache by Query-Driven Pruning
von: Xu, Yuhui, et al.
Veröffentlicht: (2024)
von: Xu, Yuhui, et al.
Veröffentlicht: (2024)
PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
von: Cai, Zefan, et al.
Veröffentlicht: (2024)
von: Cai, Zefan, et al.
Veröffentlicht: (2024)
HGMEM: Hypergraph-based Working Memory to Improve Multi-step RAG for Long-Context Complex Relational Modeling
von: Zhou, Chulun, et al.
Veröffentlicht: (2025)
von: Zhou, Chulun, et al.
Veröffentlicht: (2025)
Unsupervised Information Refinement Training of Large Language Models for Retrieval-Augmented Generation
von: Xu, Shicheng, et al.
Veröffentlicht: (2024)
von: Xu, Shicheng, et al.
Veröffentlicht: (2024)
DeltaKV: Residual-Based KV Cache Compression via Long-Range Similarity
von: Hao, Jitai, et al.
Veröffentlicht: (2026)
von: Hao, Jitai, et al.
Veröffentlicht: (2026)
Retrieval-Augmented Machine Translation with Unstructured Knowledge
von: Wang, Jiaan, et al.
Veröffentlicht: (2024)
von: Wang, Jiaan, et al.
Veröffentlicht: (2024)
TEAL: Tokenize and Embed ALL for Multi-modal Large Language Models
von: Yang, Zhen, et al.
Veröffentlicht: (2023)
von: Yang, Zhen, et al.
Veröffentlicht: (2023)
DRT: Deep Reasoning Translation via Long Chain-of-Thought
von: Wang, Jiaan, et al.
Veröffentlicht: (2024)
von: Wang, Jiaan, et al.
Veröffentlicht: (2024)
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
von: Feng, Yuan, et al.
Veröffentlicht: (2024)
von: Feng, Yuan, et al.
Veröffentlicht: (2024)
Lossless KV Cache Compression to 2%
von: Yang, Zhen, et al.
Veröffentlicht: (2024)
von: Yang, Zhen, et al.
Veröffentlicht: (2024)
G-KV: Decoding-Time KV Cache Eviction with Global Attention
von: Liao, Mengqi, et al.
Veröffentlicht: (2025)
von: Liao, Mengqi, et al.
Veröffentlicht: (2025)
LCS: A Language Converter Strategy for Zero-Shot Neural Machine Translation
von: Sun, Zengkui, et al.
Veröffentlicht: (2024)
von: Sun, Zengkui, et al.
Veröffentlicht: (2024)
Judge Q: Trainable Queries for Optimized Information Retention in KV Cache Eviction
von: Liu, Yijun, et al.
Veröffentlicht: (2025)
von: Liu, Yijun, et al.
Veröffentlicht: (2025)
RetentiveKV: State-Space Memory for Uncertainty-Aware Multimodal KV Cache Eviction
von: Liu, Sihao, et al.
Veröffentlicht: (2026)
von: Liu, Sihao, et al.
Veröffentlicht: (2026)
Training-Free Exponential Context Extension via Cascading KV Cache
von: Willette, Jeffrey, et al.
Veröffentlicht: (2024)
von: Willette, Jeffrey, et al.
Veröffentlicht: (2024)
KV Cache Steering for Controlling Frozen LLMs
von: Belitsky, Max, et al.
Veröffentlicht: (2025)
von: Belitsky, Max, et al.
Veröffentlicht: (2025)
TextBind: Multi-turn Interleaved Multimodal Instruction-following in the Wild
von: Li, Huayang, et al.
Veröffentlicht: (2023)
von: Li, Huayang, et al.
Veröffentlicht: (2023)
DivLogicEval: A Framework for Benchmarking Logical Reasoning Evaluation in Large Language Models
von: Chung, Tsz Ting, et al.
Veröffentlicht: (2025)
von: Chung, Tsz Ting, et al.
Veröffentlicht: (2025)
AhaKV: Adaptive Holistic Attention-Driven KV Cache Eviction for Efficient Inference of Large Language Models
von: Gu, Yifeng, et al.
Veröffentlicht: (2025)
von: Gu, Yifeng, et al.
Veröffentlicht: (2025)
OjaKV: Context-Aware Online Low-Rank KV Cache Compression
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
MiniCache: KV Cache Compression in Depth Dimension for Large Language Models
von: Liu, Akide, et al.
Veröffentlicht: (2024)
von: Liu, Akide, et al.
Veröffentlicht: (2024)
Continuous Autoregressive Language Models
von: Shao, Chenze, et al.
Veröffentlicht: (2025)
von: Shao, Chenze, et al.
Veröffentlicht: (2025)
General learned delegation by clones
von: Li, Darren, et al.
Veröffentlicht: (2026)
von: Li, Darren, et al.
Veröffentlicht: (2026)
Q-Filters: Leveraging QK Geometry for Efficient KV Cache Compression
von: Godey, Nathan, et al.
Veröffentlicht: (2025)
von: Godey, Nathan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A Systematic Evaluation of Large Language Models for Natural Language Generation Tasks
von: Ni, Xuanfan, et al.
Veröffentlicht: (2024) -
Marco-o1 v2: Towards Widening The Distillation Bottleneck for Reasoning Models
von: Yin, Huifeng, et al.
Veröffentlicht: (2025) -
DepthKV: Layer-Dependent KV Cache Pruning for Long-Context LLM Inference
von: Dehghanighobadi, Zahra, et al.
Veröffentlicht: (2026) -
KVzap: Fast, Adaptive, and Faithful KV Cache Pruning
von: Jegou, Simon, et al.
Veröffentlicht: (2026) -
FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference
von: Liu, Guangda, et al.
Veröffentlicht: (2025)