One Size Does Not Fit All: Token-Wise Adaptive Compression for KV Cache
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lu, Liming, Qiu, Kaixi, Zhou, Jiayu, Kai, Jushi, Zhang, Haoyan, Wang, Huanyu, Leng, Jingwen, He, Ziwei, Lin, Zhouhan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards Compressive and Scalable Recurrent Memory
von: Song, Yunchong, et al.
Veröffentlicht: (2026)
von: Song, Yunchong, et al.
Veröffentlicht: (2026)
Fourier Compressor: Frequency-Domain Visual Token Compression for Vision-Language Models
von: Wang, Huanyu, et al.
Veröffentlicht: (2025)
von: Wang, Huanyu, et al.
Veröffentlicht: (2025)
FreqKV: Key-Value Compression in Frequency Domain for Context Window Extension
von: Kai, Jushi, et al.
Veröffentlicht: (2025)
von: Kai, Jushi, et al.
Veröffentlicht: (2025)
TreeKV: Smooth Key-Value Cache Compression with Tree Structures
von: He, Ziwei, et al.
Veröffentlicht: (2025)
von: He, Ziwei, et al.
Veröffentlicht: (2025)
VQKV: High-Fidelity and High-Ratio Cache Compression via Vector-Quantization
von: Wang, Yixuan, et al.
Veröffentlicht: (2026)
von: Wang, Yixuan, et al.
Veröffentlicht: (2026)
PonderLM-3: Adaptive Token-Wise Pondering with Differentiable Masking
von: Li, He, et al.
Veröffentlicht: (2026)
von: Li, He, et al.
Veröffentlicht: (2026)
WeightedKV: Attention Scores Weighted Key-Value Cache Merging for Large Language Models
von: Yuan, Jian, et al.
Veröffentlicht: (2025)
von: Yuan, Jian, et al.
Veröffentlicht: (2025)
Meta-Soft: Leveraging Composable Meta-Tokens for Context-Preserving KV Cache Compression
von: Luo, Wei, et al.
Veröffentlicht: (2026)
von: Luo, Wei, et al.
Veröffentlicht: (2026)
AdaPonderLM: Gated Pondering Language Models with Token-Wise Adaptive Depth
von: Song, Shixiang, et al.
Veröffentlicht: (2026)
von: Song, Shixiang, et al.
Veröffentlicht: (2026)
SH2: Self-Highlighted Hesitation Helps You Decode More Truthfully
von: Kai, Jushi, et al.
Veröffentlicht: (2024)
von: Kai, Jushi, et al.
Veröffentlicht: (2024)
Leveraging Grammar Induction for Language Understanding and Generation
von: Kai, Jushi, et al.
Veröffentlicht: (2024)
von: Kai, Jushi, et al.
Veröffentlicht: (2024)
WindowKV: Task-Adaptive Group-Wise KV Cache Window Selection for Efficient LLM Inference
von: Zuo, Youhui, et al.
Veröffentlicht: (2025)
von: Zuo, Youhui, et al.
Veröffentlicht: (2025)
One Size Does Not Fit All: Architecture-Aware Adaptive Batch Scheduling with DEBA
von: Belias, François, et al.
Veröffentlicht: (2025)
von: Belias, François, et al.
Veröffentlicht: (2025)
SABlock: Semantic-Aware KV Cache Eviction with Adaptive Compression Block Size
von: Chen, Jinhan, et al.
Veröffentlicht: (2025)
von: Chen, Jinhan, et al.
Veröffentlicht: (2025)
Leak Detection Strategies: One Size Does Not Fit All
von: Jane M. Arrington
Veröffentlicht: (2026)
von: Jane M. Arrington
Veröffentlicht: (2026)
Graph-Guided Adaptive Channel Elimination for KV Cache Compression
von: Tong, Enwei, et al.
Veröffentlicht: (2026)
von: Tong, Enwei, et al.
Veröffentlicht: (2026)
Stochastic KV Routing: Enabling Adaptive Depth-Wise Cache Sharing
von: Filippova, Anastasiia, et al.
Veröffentlicht: (2026)
von: Filippova, Anastasiia, et al.
Veröffentlicht: (2026)
Uncertainty Quantification for Machine Learning: One Size Does Not Fit All
von: Hofman, Paul, et al.
Veröffentlicht: (2025)
von: Hofman, Paul, et al.
Veröffentlicht: (2025)
KVCompose: Efficient Structured KV Cache Compression with Composite Tokens
von: Akulov, Dmitry, et al.
Veröffentlicht: (2025)
von: Akulov, Dmitry, et al.
Veröffentlicht: (2025)
QAQ: Quality Adaptive Quantization for LLM KV Cache
von: Dong, Shichen, et al.
Veröffentlicht: (2024)
von: Dong, Shichen, et al.
Veröffentlicht: (2024)
One Size Does NOT Fit All: On the Importance of Physical Representations for Datalog Evaluation
von: Rassau, Nick, et al.
Veröffentlicht: (2026)
von: Rassau, Nick, et al.
Veröffentlicht: (2026)
Dosing Biologic Drugs for Patients With Obesity: One Size Does Not Fit All
von: Stephen J. Balevic, et al.
Veröffentlicht: (2026)
von: Stephen J. Balevic, et al.
Veröffentlicht: (2026)
Rethinking Stroke Prevention in Atrial Fibrillation: One Size Does not Fit All
von: Panteleimon E. Papakonstantinou, et al.
Veröffentlicht: (2025)
von: Panteleimon E. Papakonstantinou, et al.
Veröffentlicht: (2025)
ThinKV: Thought-Adaptive KV Cache Compression for Efficient Reasoning Models
von: Ramachandran, Akshat, et al.
Veröffentlicht: (2025)
von: Ramachandran, Akshat, et al.
Veröffentlicht: (2025)
KVReviver: Reversible KV Cache Compression with Sketch-Based Token Reconstruction
von: Yuan, Aomufei, et al.
Veröffentlicht: (2025)
von: Yuan, Aomufei, et al.
Veröffentlicht: (2025)
One Size Does Not Fit All: Investigating Efficacy of Perplexity in Detecting LLM-Generated Code
von: Xu, Jinwei, et al.
Veröffentlicht: (2024)
von: Xu, Jinwei, et al.
Veröffentlicht: (2024)
One Panel Does Not Fit All: Case-Adaptive Multi-Agent Deliberation for Clinical Prediction
von: Lu, Yuxing, et al.
Veröffentlicht: (2026)
von: Lu, Yuxing, et al.
Veröffentlicht: (2026)
How Much Cache Does Reasoning Need? Depth-Cache Tradeoffs in KV-Compressed Transformers
von: Wang, Xiao
Veröffentlicht: (2026)
von: Wang, Xiao
Veröffentlicht: (2026)
CaliDrop: KV Cache Compression with Calibration
von: Su, Yi, et al.
Veröffentlicht: (2025)
von: Su, Yi, et al.
Veröffentlicht: (2025)
The Pitfalls of KV Cache Compression
von: Chen, Alex, et al.
Veröffentlicht: (2025)
von: Chen, Alex, et al.
Veröffentlicht: (2025)
DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs
von: Zhou, Xiabin, et al.
Veröffentlicht: (2024)
von: Zhou, Xiabin, et al.
Veröffentlicht: (2024)
Adaptive KV-Cache Compression without Manually Setting Budget
von: Tang, Chenxia, et al.
Veröffentlicht: (2025)
von: Tang, Chenxia, et al.
Veröffentlicht: (2025)
How To Establish a Parliamentary Research Service: Does One Size Fit All?
von: Verrier, June R.
Veröffentlicht: (2000)
von: Verrier, June R.
Veröffentlicht: (2000)
The One Size Does Not Fit All Approach: Case Studies in Modeling Embedded Librarianship
von: Edford, Rachel L., et al.
Veröffentlicht: (2022)
von: Edford, Rachel L., et al.
Veröffentlicht: (2022)
One Size Does Not Fit All: Planting Calendar as an Adaptation Strategy in the Mekong Delta
von: Le Phuong‐Dung, et al.
Veröffentlicht: (2025)
von: Le Phuong‐Dung, et al.
Veröffentlicht: (2025)
When One Size Does not Fit All—Artificial Intelligence in Australian Rural Health
von: Lewis Hains, et al.
Veröffentlicht: (2025)
von: Lewis Hains, et al.
Veröffentlicht: (2025)
KV-CoRE: Benchmarking Data-Dependent Low-Rank Compressibility of KV-Caches in LLMs
von: Chen, Jian, et al.
Veröffentlicht: (2026)
von: Chen, Jian, et al.
Veröffentlicht: (2026)
Accurate KV Cache Quantization with Outlier Tokens Tracing
von: Su, Yi, et al.
Veröffentlicht: (2025)
von: Su, Yi, et al.
Veröffentlicht: (2025)
Fourier Transformer: Fast Long Range Modeling by Removing Sequence Redundancy with FFT Operator
von: He, Ziwei, et al.
Veröffentlicht: (2023)
von: He, Ziwei, et al.
Veröffentlicht: (2023)
Pyramid Cache: Layer-Adaptive KV Cache Compression with Signature-Based Cold Storage
von: Sergio dj
Veröffentlicht: (2026)
von: Sergio dj
Veröffentlicht: (2026)
Ähnliche Einträge
-
Towards Compressive and Scalable Recurrent Memory
von: Song, Yunchong, et al.
Veröffentlicht: (2026) -
Fourier Compressor: Frequency-Domain Visual Token Compression for Vision-Language Models
von: Wang, Huanyu, et al.
Veröffentlicht: (2025) -
FreqKV: Key-Value Compression in Frequency Domain for Context Window Extension
von: Kai, Jushi, et al.
Veröffentlicht: (2025) -
TreeKV: Smooth Key-Value Cache Compression with Tree Structures
von: He, Ziwei, et al.
Veröffentlicht: (2025) -
VQKV: High-Fidelity and High-Ratio Cache Compression via Vector-Quantization
von: Wang, Yixuan, et al.
Veröffentlicht: (2026)