Guardado en:
| Autores principales: | Lin, Yuping, Ding, Jiayuan, Xing, Yue, He, Pengfei, Tang, Jiliang, Mukherjee, Subhabrata |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2605.23258 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Crafting Reversible SFT Behaviors in Large Language Models
por: Lin, Yuping, et al.
Publicado: (2026)
por: Lin, Yuping, et al.
Publicado: (2026)
ManifoldKV: Training-Free KV Cache Compression via Euclidean Outlier Detection
por: Datta, Debajyoti, et al.
Publicado: (2026)
por: Datta, Debajyoti, et al.
Publicado: (2026)
EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
por: Feng, Shaoting, et al.
Publicado: (2025)
por: Feng, Shaoting, et al.
Publicado: (2025)
Beyond Token Eviction: Mixed-Dimension Budget Allocation for Efficient KV Cache Compression
por: Miao, Ruijie, et al.
Publicado: (2026)
por: Miao, Ruijie, et al.
Publicado: (2026)
In-context KV-Cache Eviction for LLMs via Attention-Gate
por: Zeng, Zihao, et al.
Publicado: (2024)
por: Zeng, Zihao, et al.
Publicado: (2024)
MomentKV: Closing the Directional Gap in KV Cache Eviction for Long-Context Inference
por: Li, Yu, et al.
Publicado: (2026)
por: Li, Yu, et al.
Publicado: (2026)
Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction
por: Bui, Ngoc, et al.
Publicado: (2026)
por: Bui, Ngoc, et al.
Publicado: (2026)
Rethinking KV Cache Eviction via a Unified Information-Theoretic Objective
por: Yang, Jiaming, et al.
Publicado: (2026)
por: Yang, Jiaming, et al.
Publicado: (2026)
LookaheadKV: Fast and Accurate KV Cache Eviction by Glimpsing into the Future without Generation
por: Ahn, Jinwoo, et al.
Publicado: (2026)
por: Ahn, Jinwoo, et al.
Publicado: (2026)
RetentiveKV: State-Space Memory for Uncertainty-Aware Multimodal KV Cache Eviction
por: Liu, Sihao, et al.
Publicado: (2026)
por: Liu, Sihao, et al.
Publicado: (2026)
Hierarchical Adaptive Eviction for KV Cache Management in Multimodal Language Models
por: Ma, Xindian, et al.
Publicado: (2026)
por: Ma, Xindian, et al.
Publicado: (2026)
LAVa: Layer-wise KV Cache Eviction with Dynamic Budget Allocation
por: Shen, Yiqun, et al.
Publicado: (2025)
por: Shen, Yiqun, et al.
Publicado: (2025)
ForesightKV: Optimizing KV Cache Eviction for Reasoning Models by Learning Long-Term Contribution
por: Dong, Zican, et al.
Publicado: (2026)
por: Dong, Zican, et al.
Publicado: (2026)
When Does Value-Aware KV Eviction Help? A Fixed-Contract Diagnostic for Non-Monotone Cache Compression
por: Zhang, Ruijie, et al.
Publicado: (2026)
por: Zhang, Ruijie, et al.
Publicado: (2026)
CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM
por: Li, Yubo, et al.
Publicado: (2026)
por: Li, Yubo, et al.
Publicado: (2026)
CAOTE: KV Cache Selection for LLMs via Attention Output Error-Based Token Eviction
por: Goel, Raghavv, et al.
Publicado: (2025)
por: Goel, Raghavv, et al.
Publicado: (2025)
RDKV: Rate-Distortion Bit Allocation for Joint Eviction and Quantization of the KV Cache
por: Zhang, Junkai, et al.
Publicado: (2026)
por: Zhang, Junkai, et al.
Publicado: (2026)
PagedEviction: Structured Block-wise KV Cache Pruning for Efficient Large Language Model Inference
por: Chitty-Venkata, Krishna Teja, et al.
Publicado: (2025)
por: Chitty-Venkata, Krishna Teja, et al.
Publicado: (2025)
LKV: End-to-End Learning of Head-wise Budgets and Token Selection for LLM KV Cache Eviction
por: Zhou, Enshuai, et al.
Publicado: (2026)
por: Zhou, Enshuai, et al.
Publicado: (2026)
Make LLMs better zero-shot reasoners: Structure-orientated autonomous reasoning
por: He, Pengfei, et al.
Publicado: (2024)
por: He, Pengfei, et al.
Publicado: (2024)
MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context Inference
por: Li, Kunxi, et al.
Publicado: (2025)
por: Li, Kunxi, et al.
Publicado: (2025)
The Pitfalls of KV Cache Compression
por: Chen, Alex, et al.
Publicado: (2025)
por: Chen, Alex, et al.
Publicado: (2025)
RazorAttention: Efficient KV Cache Compression Through Retrieval Heads
por: Tang, Hanlin, et al.
Publicado: (2024)
por: Tang, Hanlin, et al.
Publicado: (2024)
Training Transformers for KV Cache Compressibility
por: Gelberg, Yoav, et al.
Publicado: (2026)
por: Gelberg, Yoav, et al.
Publicado: (2026)
Multi-Faceted Studies on Data Poisoning can Advance LLM Development
por: He, Pengfei, et al.
Publicado: (2025)
por: He, Pengfei, et al.
Publicado: (2025)
Superiority of Multi-Head Attention in In-Context Linear Regression
por: Cui, Yingqian, et al.
Publicado: (2024)
por: Cui, Yingqian, et al.
Publicado: (2024)
Palu: Compressing KV-Cache with Low-Rank Projection
por: Chang, Chi-Chih, et al.
Publicado: (2024)
por: Chang, Chi-Chih, et al.
Publicado: (2024)
KVSculpt: KV Cache Compression as Distillation
por: Jiang, Bo, et al.
Publicado: (2026)
por: Jiang, Bo, et al.
Publicado: (2026)
KV-CAR: KV Cache Compression using Autoencoders and KV Reuse in Large Language Models
por: Roy, Sourjya, et al.
Publicado: (2025)
por: Roy, Sourjya, et al.
Publicado: (2025)
AttentionPredictor: Temporal Patterns Matter for KV Cache Compression
por: Yang, Qingyue, et al.
Publicado: (2025)
por: Yang, Qingyue, et al.
Publicado: (2025)
LazyEviction: Lagged KV Eviction with Attention Pattern Observation for Efficient Long Reasoning
por: Zhang, Haoyue, et al.
Publicado: (2025)
por: Zhang, Haoyue, et al.
Publicado: (2025)
MiniCache: KV Cache Compression in Depth Dimension for Large Language Models
por: Liu, Akide, et al.
Publicado: (2024)
por: Liu, Akide, et al.
Publicado: (2024)
ThinKV: Thought-Adaptive KV Cache Compression for Efficient Reasoning Models
por: Ramachandran, Akshat, et al.
Publicado: (2025)
por: Ramachandran, Akshat, et al.
Publicado: (2025)
CommonKV: Compressing KV Cache with Cross-layer Parameter Sharing
por: Wang, Yixuan, et al.
Publicado: (2025)
por: Wang, Yixuan, et al.
Publicado: (2025)
xKV: Cross-Layer KV-Cache Compression via Aligned Singular Vector Extraction
por: Chang, Chi-Chih, et al.
Publicado: (2025)
por: Chang, Chi-Chih, et al.
Publicado: (2025)
EvolKV: Evolutionary KV Cache Compression for LLM Inference
por: Yu, Bohan, et al.
Publicado: (2025)
por: Yu, Bohan, et al.
Publicado: (2025)
A Theoretical Understanding of Chain-of-Thought: Coherent Reasoning and Error-Aware Demonstration
por: Cui, Yingqian, et al.
Publicado: (2024)
por: Cui, Yingqian, et al.
Publicado: (2024)
Towards the Effect of Examples on In-Context Learning: A Theoretical Case Study
por: He, Pengfei, et al.
Publicado: (2024)
por: He, Pengfei, et al.
Publicado: (2024)
HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing
por: Liu, Minghui, et al.
Publicado: (2024)
por: Liu, Minghui, et al.
Publicado: (2024)
LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction for Efficient Long-Context Inference
por: Wang, Guangtao, et al.
Publicado: (2025)
por: Wang, Guangtao, et al.
Publicado: (2025)
Ejemplares similares
-
Crafting Reversible SFT Behaviors in Large Language Models
por: Lin, Yuping, et al.
Publicado: (2026) -
ManifoldKV: Training-Free KV Cache Compression via Euclidean Outlier Detection
por: Datta, Debajyoti, et al.
Publicado: (2026) -
EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
por: Feng, Shaoting, et al.
Publicado: (2025) -
Beyond Token Eviction: Mixed-Dimension Budget Allocation for Efficient KV Cache Compression
por: Miao, Ruijie, et al.
Publicado: (2026) -
In-context KV-Cache Eviction for LLMs via Attention-Gate
por: Zeng, Zihao, et al.
Publicado: (2024)