AhaKV: Adaptive Holistic Attention-Driven KV Cache Eviction for Efficient Inference of Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gu, Yifeng, Jiang, Zicong, Jin, Jianxiu, Guo, Kailing, Zhang, Ziyang, Xu, Xiangmin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
von: Feng, Yuan, et al.
Veröffentlicht: (2024)
von: Feng, Yuan, et al.
Veröffentlicht: (2024)
MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context Inference
von: Li, Kunxi, et al.
Veröffentlicht: (2025)
von: Li, Kunxi, et al.
Veröffentlicht: (2025)
Taming the Fragility of KV Cache Eviction in LLM Inference
von: Feng, Yuan, et al.
Veröffentlicht: (2025)
von: Feng, Yuan, et al.
Veröffentlicht: (2025)
G-KV: Decoding-Time KV Cache Eviction with Global Attention
von: Liao, Mengqi, et al.
Veröffentlicht: (2025)
von: Liao, Mengqi, et al.
Veröffentlicht: (2025)
CriticalKV: Optimizing KV Cache Eviction from an Output Perturbation Perspective
von: Feng, Yuan, et al.
Veröffentlicht: (2025)
von: Feng, Yuan, et al.
Veröffentlicht: (2025)
GraphKV: Breaking the Static Selection Paradigm with Graph-Based KV Cache Eviction
von: Li, Xuelin, et al.
Veröffentlicht: (2025)
von: Li, Xuelin, et al.
Veröffentlicht: (2025)
CAKE: Cascading and Adaptive KV Cache Eviction with Layer Preferences
von: Qin, Ziran, et al.
Veröffentlicht: (2025)
von: Qin, Ziran, et al.
Veröffentlicht: (2025)
In-context KV-Cache Eviction for LLMs via Attention-Gate
von: Zeng, Zihao, et al.
Veröffentlicht: (2024)
von: Zeng, Zihao, et al.
Veröffentlicht: (2024)
KVDrive: A Holistic Multi-Tier KV Cache Management System for Long-Context LLM Inference
von: Lin, Jian, et al.
Veröffentlicht: (2026)
von: Lin, Jian, et al.
Veröffentlicht: (2026)
IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM Inference
von: Yang, Xintong, et al.
Veröffentlicht: (2026)
von: Yang, Xintong, et al.
Veröffentlicht: (2026)
SABlock: Semantic-Aware KV Cache Eviction with Adaptive Compression Block Size
von: Chen, Jinhan, et al.
Veröffentlicht: (2025)
von: Chen, Jinhan, et al.
Veröffentlicht: (2025)
Reformulating KV Cache Eviction Problem for Long-Context LLM Inference
von: Mai, Tho, et al.
Veröffentlicht: (2026)
von: Mai, Tho, et al.
Veröffentlicht: (2026)
RetentiveKV: State-Space Memory for Uncertainty-Aware Multimodal KV Cache Eviction
von: Liu, Sihao, et al.
Veröffentlicht: (2026)
von: Liu, Sihao, et al.
Veröffentlicht: (2026)
LazyEviction: Lagged KV Eviction with Attention Pattern Observation for Efficient Long Reasoning
von: Zhang, Haoyue, et al.
Veröffentlicht: (2025)
von: Zhang, Haoyue, et al.
Veröffentlicht: (2025)
LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction for Efficient Long-Context Inference
von: Wang, Guangtao, et al.
Veröffentlicht: (2025)
von: Wang, Guangtao, et al.
Veröffentlicht: (2025)
AudioKV: KV Cache Eviction in Efficient Large Audio Language Models
von: Wang, Yuxuan, et al.
Veröffentlicht: (2026)
von: Wang, Yuxuan, et al.
Veröffentlicht: (2026)
NACL: A General and Effective KV Cache Eviction Framework for LLMs at Inference Time
von: Chen, Yilong, et al.
Veröffentlicht: (2024)
von: Chen, Yilong, et al.
Veröffentlicht: (2024)
ForesightKV: Optimizing KV Cache Eviction for Reasoning Models by Learning Long-Term Contribution
von: Dong, Zican, et al.
Veröffentlicht: (2026)
von: Dong, Zican, et al.
Veröffentlicht: (2026)
FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference
von: Liu, Guangda, et al.
Veröffentlicht: (2025)
von: Liu, Guangda, et al.
Veröffentlicht: (2025)
RobustKV: Defending Large Language Models against Jailbreak Attacks via KV Eviction
von: Jiang, Tanqiu, et al.
Veröffentlicht: (2024)
von: Jiang, Tanqiu, et al.
Veröffentlicht: (2024)
Layer-Condensed KV Cache for Efficient Inference of Large Language Models
von: Wu, Haoyi, et al.
Veröffentlicht: (2024)
von: Wu, Haoyi, et al.
Veröffentlicht: (2024)
Lookahead Q-Cache: Achieving More Consistent KV Cache Eviction via Pseudo Query
von: Wang, Yixuan, et al.
Veröffentlicht: (2025)
von: Wang, Yixuan, et al.
Veröffentlicht: (2025)
Judge Q: Trainable Queries for Optimized Information Retention in KV Cache Eviction
von: Liu, Yijun, et al.
Veröffentlicht: (2025)
von: Liu, Yijun, et al.
Veröffentlicht: (2025)
WindowKV: Task-Adaptive Group-Wise KV Cache Window Selection for Efficient LLM Inference
von: Zuo, Youhui, et al.
Veröffentlicht: (2025)
von: Zuo, Youhui, et al.
Veröffentlicht: (2025)
ReST-KV: Robust KV Cache Eviction with Layer-wise Output Reconstruction and Spatial-Temporal Smoothing
von: An, Yongqi, et al.
Veröffentlicht: (2026)
von: An, Yongqi, et al.
Veröffentlicht: (2026)
Fast KVzip: Efficient and Accurate LLM Inference with Gated KV Eviction
von: Kim, Jang-Hyun, et al.
Veröffentlicht: (2026)
von: Kim, Jang-Hyun, et al.
Veröffentlicht: (2026)
CAOTE: KV Cache Selection for LLMs via Attention Output Error-Based Token Eviction
von: Goel, Raghavv, et al.
Veröffentlicht: (2025)
von: Goel, Raghavv, et al.
Veröffentlicht: (2025)
SpindleKV: A Novel KV Cache Reduction Method Balancing Both Shallow and Deep Layers
von: Tang, Zicong, et al.
Veröffentlicht: (2025)
von: Tang, Zicong, et al.
Veröffentlicht: (2025)
ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference
von: Liu, Xiang, et al.
Veröffentlicht: (2025)
von: Liu, Xiang, et al.
Veröffentlicht: (2025)
KVSculpt: KV Cache Compression as Distillation
von: Jiang, Bo, et al.
Veröffentlicht: (2026)
von: Jiang, Bo, et al.
Veröffentlicht: (2026)
SentenceKV: Efficient LLM Inference via Sentence-Level Semantic KV Caching
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
von: Tian, Yuxuan, et al.
Veröffentlicht: (2025)
von: Tian, Yuxuan, et al.
Veröffentlicht: (2025)
EvolKV: Evolutionary KV Cache Compression for LLM Inference
von: Yu, Bohan, et al.
Veröffentlicht: (2025)
von: Yu, Bohan, et al.
Veröffentlicht: (2025)
MPCache: MPC-Friendly KV Cache Eviction for Efficient Private LLM Inference
von: Zeng, Wenxuan, et al.
Veröffentlicht: (2025)
von: Zeng, Wenxuan, et al.
Veröffentlicht: (2025)
DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs
von: Zhou, Xiabin, et al.
Veröffentlicht: (2024)
von: Zhou, Xiabin, et al.
Veröffentlicht: (2024)
LKV: End-to-End Learning of Head-wise Budgets and Token Selection for LLM KV Cache Eviction
von: Zhou, Enshuai, et al.
Veröffentlicht: (2026)
von: Zhou, Enshuai, et al.
Veröffentlicht: (2026)
HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing
von: Liu, Minghui, et al.
Veröffentlicht: (2024)
von: Liu, Minghui, et al.
Veröffentlicht: (2024)
MomentKV: Closing the Directional Gap in KV Cache Eviction for Long-Context Inference
von: Li, Yu, et al.
Veröffentlicht: (2026)
von: Li, Yu, et al.
Veröffentlicht: (2026)
Beyond KV Caching: Shared Attention for Efficient LLMs
von: Liao, Bingli, et al.
Veröffentlicht: (2024)
von: Liao, Bingli, et al.
Veröffentlicht: (2024)
WeightedKV: Attention Scores Weighted Key-Value Cache Merging for Large Language Models
von: Yuan, Jian, et al.
Veröffentlicht: (2025)
von: Yuan, Jian, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
von: Feng, Yuan, et al.
Veröffentlicht: (2024) -
MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context Inference
von: Li, Kunxi, et al.
Veröffentlicht: (2025) -
Taming the Fragility of KV Cache Eviction in LLM Inference
von: Feng, Yuan, et al.
Veröffentlicht: (2025) -
G-KV: Decoding-Time KV Cache Eviction with Global Attention
von: Liao, Mengqi, et al.
Veröffentlicht: (2025) -
CriticalKV: Optimizing KV Cache Eviction from an Output Perturbation Perspective
von: Feng, Yuan, et al.
Veröffentlicht: (2025)