Understanding the Physics of Key-Value Cache Compression for LLMs through Attention Dynamics
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ananthanarayanan, Samhruth, Sengupta, Ayan, Chakraborty, Tanmoy |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Value-Guided KV Compression for LLMs via Approximated CUR Decomposition
von: Sengupta, Ayan, et al.
Veröffentlicht: (2025)
von: Sengupta, Ayan, et al.
Veröffentlicht: (2025)
Compression Laws for Large Language Models
von: Sengupta, Ayan, et al.
Veröffentlicht: (2025)
von: Sengupta, Ayan, et al.
Veröffentlicht: (2025)
You Only Prune Once: Designing Calibration-Free Model Compression With Policy Learning
von: Sengupta, Ayan, et al.
Veröffentlicht: (2025)
von: Sengupta, Ayan, et al.
Veröffentlicht: (2025)
Position: Enough of Scaling LLMs! Lets Focus on Downscaling
von: Goel, Yash, et al.
Veröffentlicht: (2025)
von: Goel, Yash, et al.
Veröffentlicht: (2025)
The Art of Scaling Test-Time Compute for Large Language Models
von: Agarwal, Aradhye, et al.
Veröffentlicht: (2025)
von: Agarwal, Aradhye, et al.
Veröffentlicht: (2025)
First Finish Search: Efficient Test-Time Scaling in Large Language Models
von: Agarwal, Aradhye, et al.
Veröffentlicht: (2025)
von: Agarwal, Aradhye, et al.
Veröffentlicht: (2025)
How to Upscale Neural Networks with Scaling Law? A Survey and Practical Guidelines
von: Sengupta, Ayan, et al.
Veröffentlicht: (2025)
von: Sengupta, Ayan, et al.
Veröffentlicht: (2025)
On the Generalization vs Fidelity Paradox in Knowledge Distillation
von: Ramesh, Suhas Kamasetty, et al.
Veröffentlicht: (2025)
von: Ramesh, Suhas Kamasetty, et al.
Veröffentlicht: (2025)
Persona-aware Generative Model for Code-mixed Language
von: Sengupta, Ayan, et al.
Veröffentlicht: (2023)
von: Sengupta, Ayan, et al.
Veröffentlicht: (2023)
From Images to Words: Efficient Cross-Modal Knowledge Distillation to Language Models from Black-box Teachers
von: Sengupta, Ayan, et al.
Veröffentlicht: (2026)
von: Sengupta, Ayan, et al.
Veröffentlicht: (2026)
Step-by-Step Unmasking for Parameter-Efficient Fine-tuning of Large Language Models
von: Agarwal, Aradhye, et al.
Veröffentlicht: (2024)
von: Agarwal, Aradhye, et al.
Veröffentlicht: (2024)
SpeCache: Speculative Key-Value Caching for Efficient Generation of LLMs
von: Jie, Shibo, et al.
Veröffentlicht: (2025)
von: Jie, Shibo, et al.
Veröffentlicht: (2025)
Key, Value, Compress: A Systematic Exploration of KV Cache Compression Techniques
von: Javidnia, Neusha, et al.
Veröffentlicht: (2025)
von: Javidnia, Neusha, et al.
Veröffentlicht: (2025)
SONIC: Segmented Optimized Nexus for Information Compression in Key-Value Caching
von: Chen, Hong, et al.
Veröffentlicht: (2026)
von: Chen, Hong, et al.
Veröffentlicht: (2026)
TreeKV: Smooth Key-Value Cache Compression with Tree Structures
von: He, Ziwei, et al.
Veröffentlicht: (2025)
von: He, Ziwei, et al.
Veröffentlicht: (2025)
SCOPE: Optimizing Key-Value Cache Compression in Long-context Generation
von: Wu, Jialong, et al.
Veröffentlicht: (2024)
von: Wu, Jialong, et al.
Veröffentlicht: (2024)
Understanding the Effects of Domain Finetuning on LLMs
von: Tanwar, Eshaan, et al.
Veröffentlicht: (2025)
von: Tanwar, Eshaan, et al.
Veröffentlicht: (2025)
Robust and Efficient Fine-tuning of LLMs with Bayesian Reparameterization of Low-Rank Adaptation
von: Sengupta, Ayan, et al.
Veröffentlicht: (2024)
von: Sengupta, Ayan, et al.
Veröffentlicht: (2024)
Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
von: Brandon, William, et al.
Veröffentlicht: (2024)
von: Brandon, William, et al.
Veröffentlicht: (2024)
Siren -- Advancing Cybersecurity through Deception and Adaptive Analysis
von: Ananthanarayanan, Samhruth, et al.
Veröffentlicht: (2024)
von: Ananthanarayanan, Samhruth, et al.
Veröffentlicht: (2024)
WeightedKV: Attention Scores Weighted Key-Value Cache Merging for Large Language Models
von: Yuan, Jian, et al.
Veröffentlicht: (2025)
von: Yuan, Jian, et al.
Veröffentlicht: (2025)
Trellis: Learning to Compress Key-Value Memory in Attention Models
von: Karami, Mahdi, et al.
Veröffentlicht: (2025)
von: Karami, Mahdi, et al.
Veröffentlicht: (2025)
KVSink: Understanding and Enhancing the Preservation of Attention Sinks in KV Cache Quantization for LLMs
von: Su, Zunhai, et al.
Veröffentlicht: (2025)
von: Su, Zunhai, et al.
Veröffentlicht: (2025)
Multilingual LLMs Struggle to Link Orthography and Semantics in Bilingual Word Processing
von: Tanwar, Eshaan, et al.
Veröffentlicht: (2025)
von: Tanwar, Eshaan, et al.
Veröffentlicht: (2025)
Temporal Referential Consistency: Do LLMs Favor Sequences Over Absolute Time References?
von: Bajpai, Ashutosh, et al.
Veröffentlicht: (2025)
von: Bajpai, Ashutosh, et al.
Veröffentlicht: (2025)
Massive Values in Self-Attention Modules are the Key to Contextual Knowledge Understanding
von: Jin, Mingyu, et al.
Veröffentlicht: (2025)
von: Jin, Mingyu, et al.
Veröffentlicht: (2025)
Learning to Evict from Key-Value Cache
von: Moschella, Luca, et al.
Veröffentlicht: (2026)
von: Moschella, Luca, et al.
Veröffentlicht: (2026)
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs
von: Yang, Dongquan, et al.
Veröffentlicht: (2025)
von: Yang, Dongquan, et al.
Veröffentlicht: (2025)
DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs
von: Zhou, Xiabin, et al.
Veröffentlicht: (2024)
von: Zhou, Xiabin, et al.
Veröffentlicht: (2024)
Counterspeech the ultimate shield! Multi-Conditioned Counterspeech Generation through Attributed Prefix Learning
von: Kumar, Aswini, et al.
Veröffentlicht: (2025)
von: Kumar, Aswini, et al.
Veröffentlicht: (2025)
FlowKV: Enhancing Multi-Turn Conversational Coherence in LLMs via Isolated Key-Value Cache Management
von: Liu, Xiang, et al.
Veröffentlicht: (2025)
von: Liu, Xiang, et al.
Veröffentlicht: (2025)
KV-Compress: Paged KV-Cache Compression with Variable Compression Rates per Attention Head
von: Rehg, Isaac
Veröffentlicht: (2024)
von: Rehg, Isaac
Veröffentlicht: (2024)
Exposing Long-Tail Safety Failures in Large Language Models through Efficient Diverse Response Sampling
von: Hajra, Suvadeep, et al.
Veröffentlicht: (2026)
von: Hajra, Suvadeep, et al.
Veröffentlicht: (2026)
SVDq: 1.25-bit and 410x Key Cache Compression for LLM Attention
von: Yankun, Hong, et al.
Veröffentlicht: (2025)
von: Yankun, Hong, et al.
Veröffentlicht: (2025)
Enhancing IoT based Plant Health Monitoring through Advanced Human Plant Interaction using Large Language Models and Mobile Applications
von: Agarwal, Kriti, et al.
Veröffentlicht: (2024)
von: Agarwal, Kriti, et al.
Veröffentlicht: (2024)
Homogeneous Keys, Heterogeneous Values: Exploiting Local KV Cache Asymmetry for Long-Context LLMs
von: Cui, Wanyun, et al.
Veröffentlicht: (2025)
von: Cui, Wanyun, et al.
Veröffentlicht: (2025)
Can LLMs replace Neil deGrasse Tyson? Evaluating the Reliability of LLMs as Science Communicators
von: Bajpai, Prasoon, et al.
Veröffentlicht: (2024)
von: Bajpai, Prasoon, et al.
Veröffentlicht: (2024)
Multilingual LLMs Inherently Reward In-Language Time-Sensitive Semantic Alignment for Low-Resource Languages
von: Bajpai, Ashutosh, et al.
Veröffentlicht: (2024)
von: Bajpai, Ashutosh, et al.
Veröffentlicht: (2024)
Unshackling Context Length: An Efficient Selective Attention Approach through Query-Key Compression
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
Multilingual Test-Time Scaling via Initial Thought Transfer
von: Bajpai, Prasoon, et al.
Veröffentlicht: (2025)
von: Bajpai, Prasoon, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Value-Guided KV Compression for LLMs via Approximated CUR Decomposition
von: Sengupta, Ayan, et al.
Veröffentlicht: (2025) -
Compression Laws for Large Language Models
von: Sengupta, Ayan, et al.
Veröffentlicht: (2025) -
You Only Prune Once: Designing Calibration-Free Model Compression With Policy Learning
von: Sengupta, Ayan, et al.
Veröffentlicht: (2025) -
Position: Enough of Scaling LLMs! Lets Focus on Downscaling
von: Goel, Yash, et al.
Veröffentlicht: (2025) -
The Art of Scaling Test-Time Compute for Large Language Models
von: Agarwal, Aradhye, et al.
Veröffentlicht: (2025)