LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction for Efficient Long-Context Inference
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Guangtao, Upasani, Shubhangi, Wu, Chen, Gandhi, Darshan, Li, Jonathan, Hu, Changran, Li, Bo, Thakker, Urmish |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
The Limits of Long-Context Reasoning in Automated Bug Fixing
por: Raju, Ravi, et al.
Publicado: (2026)
por: Raju, Ravi, et al.
Publicado: (2026)
Test-Time Adaptation via Many-Shot Prompting: Benefits, Limits, and Pitfalls
por: Upasani, Shubhangi, et al.
Publicado: (2026)
por: Upasani, Shubhangi, et al.
Publicado: (2026)
Cross-Family Speculative Prefill: Training-Free Long-Context Compression with Small Draft Models
por: Upasani, Shubhangi, et al.
Publicado: (2026)
por: Upasani, Shubhangi, et al.
Publicado: (2026)
MomentKV: Closing the Directional Gap in KV Cache Eviction for Long-Context Inference
por: Li, Yu, et al.
Publicado: (2026)
por: Li, Yu, et al.
Publicado: (2026)
Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models
por: Zhang, Qizheng, et al.
Publicado: (2025)
por: Zhang, Qizheng, et al.
Publicado: (2025)
MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context Inference
por: Li, Kunxi, et al.
Publicado: (2025)
por: Li, Kunxi, et al.
Publicado: (2025)
Reformulating KV Cache Eviction Problem for Long-Context LLM Inference
por: Mai, Tho, et al.
Publicado: (2026)
por: Mai, Tho, et al.
Publicado: (2026)
In-context KV-Cache Eviction for LLMs via Attention-Gate
por: Zeng, Zihao, et al.
Publicado: (2024)
por: Zeng, Zihao, et al.
Publicado: (2024)
IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM Inference
por: Yang, Xintong, et al.
Publicado: (2026)
por: Yang, Xintong, et al.
Publicado: (2026)
AhaKV: Adaptive Holistic Attention-Driven KV Cache Eviction for Efficient Inference of Large Language Models
por: Gu, Yifeng, et al.
Publicado: (2025)
por: Gu, Yifeng, et al.
Publicado: (2025)
MPCache: MPC-Friendly KV Cache Eviction for Efficient Private LLM Inference
por: Zeng, Wenxuan, et al.
Publicado: (2025)
por: Zeng, Wenxuan, et al.
Publicado: (2025)
ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference
por: Liu, Xiang, et al.
Publicado: (2025)
por: Liu, Xiang, et al.
Publicado: (2025)
G-KV: Decoding-Time KV Cache Eviction with Global Attention
por: Liao, Mengqi, et al.
Publicado: (2025)
por: Liao, Mengqi, et al.
Publicado: (2025)
Taming the Fragility of KV Cache Eviction in LLM Inference
por: Feng, Yuan, et al.
Publicado: (2025)
por: Feng, Yuan, et al.
Publicado: (2025)
SubgoalXL: Subgoal-based Expert Learning for Theorem Proving
por: Zhao, Xueliang, et al.
Publicado: (2024)
por: Zhao, Xueliang, et al.
Publicado: (2024)
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
por: Feng, Yuan, et al.
Publicado: (2024)
por: Feng, Yuan, et al.
Publicado: (2024)
LazyEviction: Lagged KV Eviction with Attention Pattern Observation for Efficient Long Reasoning
por: Zhang, Haoyue, et al.
Publicado: (2025)
por: Zhang, Haoyue, et al.
Publicado: (2025)
HillInfer: Efficient Long-Context LLM Inference on the Edge with Hierarchical KV Eviction using SmartSSD
por: Sun, He, et al.
Publicado: (2026)
por: Sun, He, et al.
Publicado: (2026)
NACL: A General and Effective KV Cache Eviction Framework for LLMs at Inference Time
por: Chen, Yilong, et al.
Publicado: (2024)
por: Chen, Yilong, et al.
Publicado: (2024)
CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM
por: Li, Yubo, et al.
Publicado: (2026)
por: Li, Yubo, et al.
Publicado: (2026)
Training Domain Draft Models for Speculative Decoding: Best Practices and Insights
por: Hong, Fenglu, et al.
Publicado: (2025)
por: Hong, Fenglu, et al.
Publicado: (2025)
ForesightKV: Optimizing KV Cache Eviction for Reasoning Models by Learning Long-Term Contribution
por: Dong, Zican, et al.
Publicado: (2026)
por: Dong, Zican, et al.
Publicado: (2026)
KeyDiff: Key Similarity-Based KV Cache Eviction for Long-Context LLM Inference in Resource-Constrained Environments
por: Park, Junyoung, et al.
Publicado: (2025)
por: Park, Junyoung, et al.
Publicado: (2025)
CAOTE: KV Cache Selection for LLMs via Attention Output Error-Based Token Eviction
por: Goel, Raghavv, et al.
Publicado: (2025)
por: Goel, Raghavv, et al.
Publicado: (2025)
Efficient Long-Context LLM Inference via KV Cache Clustering
por: Hu, Jie, et al.
Publicado: (2025)
por: Hu, Jie, et al.
Publicado: (2025)
Constructing Domain-Specific Evaluation Sets for LLM-as-a-judge
por: Raju, Ravi, et al.
Publicado: (2024)
por: Raju, Ravi, et al.
Publicado: (2024)
AudioKV: KV Cache Eviction in Efficient Large Audio Language Models
por: Wang, Yuxuan, et al.
Publicado: (2026)
por: Wang, Yuxuan, et al.
Publicado: (2026)
Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction
por: Bui, Ngoc, et al.
Publicado: (2026)
por: Bui, Ngoc, et al.
Publicado: (2026)
SnapStream: Efficient Long Sequence Decoding on Dataflow Accelerators
por: Li, Jonathan, et al.
Publicado: (2025)
por: Li, Jonathan, et al.
Publicado: (2025)
EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
por: Feng, Shaoting, et al.
Publicado: (2025)
por: Feng, Shaoting, et al.
Publicado: (2025)
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity
por: Ma, Da, et al.
Publicado: (2024)
por: Ma, Da, et al.
Publicado: (2024)
GraphKV: Breaking the Static Selection Paradigm with Graph-Based KV Cache Eviction
por: Li, Xuelin, et al.
Publicado: (2025)
por: Li, Xuelin, et al.
Publicado: (2025)
CAKE: Cascading and Adaptive KV Cache Eviction with Layer Preferences
por: Qin, Ziran, et al.
Publicado: (2025)
por: Qin, Ziran, et al.
Publicado: (2025)
PagedEviction: Structured Block-wise KV Cache Pruning for Efficient Large Language Model Inference
por: Chitty-Venkata, Krishna Teja, et al.
Publicado: (2025)
por: Chitty-Venkata, Krishna Teja, et al.
Publicado: (2025)
ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference
por: Sun, Hanshi, et al.
Publicado: (2024)
por: Sun, Hanshi, et al.
Publicado: (2024)
Beyond Token Eviction: Mixed-Dimension Budget Allocation for Efficient KV Cache Compression
por: Miao, Ruijie, et al.
Publicado: (2026)
por: Miao, Ruijie, et al.
Publicado: (2026)
DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs
por: Zhou, Xiabin, et al.
Publicado: (2024)
por: Zhou, Xiabin, et al.
Publicado: (2024)
CaliDrop: KV Cache Compression with Calibration
por: Su, Yi, et al.
Publicado: (2025)
por: Su, Yi, et al.
Publicado: (2025)
MEDA: Dynamic KV Cache Allocation for Efficient Multimodal Long-Context Inference
por: Wan, Zhongwei, et al.
Publicado: (2025)
por: Wan, Zhongwei, et al.
Publicado: (2025)
SambaLingo: Teaching Large Language Models New Languages
por: Csaki, Zoltan, et al.
Publicado: (2024)
por: Csaki, Zoltan, et al.
Publicado: (2024)
Ejemplares similares
-
The Limits of Long-Context Reasoning in Automated Bug Fixing
por: Raju, Ravi, et al.
Publicado: (2026) -
Test-Time Adaptation via Many-Shot Prompting: Benefits, Limits, and Pitfalls
por: Upasani, Shubhangi, et al.
Publicado: (2026) -
Cross-Family Speculative Prefill: Training-Free Long-Context Compression with Small Draft Models
por: Upasani, Shubhangi, et al.
Publicado: (2026) -
MomentKV: Closing the Directional Gap in KV Cache Eviction for Long-Context Inference
por: Li, Yu, et al.
Publicado: (2026) -
Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models
por: Zhang, Qizheng, et al.
Publicado: (2025)