Value-Guided KV Compression for LLMs via Approximated CUR Decomposition
Fuente:
arXiv
Saved in:
| Main Authors: | Sengupta, Ayan, Chaudhary, Siddhant, Chakraborty, Tanmoy |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Compression Laws for Large Language Models
by: Sengupta, Ayan, et al.
Published: (2025)
by: Sengupta, Ayan, et al.
Published: (2025)
You Only Prune Once: Designing Calibration-Free Model Compression With Policy Learning
by: Sengupta, Ayan, et al.
Published: (2025)
by: Sengupta, Ayan, et al.
Published: (2025)
Understanding the Physics of Key-Value Cache Compression for LLMs through Attention Dynamics
by: Ananthanarayanan, Samhruth, et al.
Published: (2026)
by: Ananthanarayanan, Samhruth, et al.
Published: (2026)
Position: Enough of Scaling LLMs! Lets Focus on Downscaling
by: Goel, Yash, et al.
Published: (2025)
by: Goel, Yash, et al.
Published: (2025)
The Art of Scaling Test-Time Compute for Large Language Models
by: Agarwal, Aradhye, et al.
Published: (2025)
by: Agarwal, Aradhye, et al.
Published: (2025)
First Finish Search: Efficient Test-Time Scaling in Large Language Models
by: Agarwal, Aradhye, et al.
Published: (2025)
by: Agarwal, Aradhye, et al.
Published: (2025)
How to Upscale Neural Networks with Scaling Law? A Survey and Practical Guidelines
by: Sengupta, Ayan, et al.
Published: (2025)
by: Sengupta, Ayan, et al.
Published: (2025)
On the Generalization vs Fidelity Paradox in Knowledge Distillation
by: Ramesh, Suhas Kamasetty, et al.
Published: (2025)
by: Ramesh, Suhas Kamasetty, et al.
Published: (2025)
Persona-aware Generative Model for Code-mixed Language
by: Sengupta, Ayan, et al.
Published: (2023)
by: Sengupta, Ayan, et al.
Published: (2023)
From Images to Words: Efficient Cross-Modal Knowledge Distillation to Language Models from Black-box Teachers
by: Sengupta, Ayan, et al.
Published: (2026)
by: Sengupta, Ayan, et al.
Published: (2026)
Step-by-Step Unmasking for Parameter-Efficient Fine-tuning of Large Language Models
by: Agarwal, Aradhye, et al.
Published: (2024)
by: Agarwal, Aradhye, et al.
Published: (2024)
MemeMQA: Multimodal Question Answering for Memes via Rationale-Based Inferencing
by: Agarwal, Siddhant, et al.
Published: (2024)
by: Agarwal, Siddhant, et al.
Published: (2024)
Robust and Efficient Fine-tuning of LLMs with Bayesian Reparameterization of Low-Rank Adaptation
by: Sengupta, Ayan, et al.
Published: (2024)
by: Sengupta, Ayan, et al.
Published: (2024)
Key, Value, Compress: A Systematic Exploration of KV Cache Compression Techniques
by: Javidnia, Neusha, et al.
Published: (2025)
by: Javidnia, Neusha, et al.
Published: (2025)
DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs
by: Zhou, Xiabin, et al.
Published: (2024)
by: Zhou, Xiabin, et al.
Published: (2024)
Multilingual Test-Time Scaling via Initial Thought Transfer
by: Bajpai, Prasoon, et al.
Published: (2025)
by: Bajpai, Prasoon, et al.
Published: (2025)
Multilingual LLMs Struggle to Link Orthography and Semantics in Bilingual Word Processing
by: Tanwar, Eshaan, et al.
Published: (2025)
by: Tanwar, Eshaan, et al.
Published: (2025)
Temporal Referential Consistency: Do LLMs Favor Sequences Over Absolute Time References?
by: Bajpai, Ashutosh, et al.
Published: (2025)
by: Bajpai, Ashutosh, et al.
Published: (2025)
KV-CoRE: Benchmarking Data-Dependent Low-Rank Compressibility of KV-Caches in LLMs
by: Chen, Jian, et al.
Published: (2026)
by: Chen, Jian, et al.
Published: (2026)
Beyond Homogeneous Attention: Memory-Efficient LLMs via Fourier-Approximated KV Cache
by: Liu, Xiaoran, et al.
Published: (2025)
by: Liu, Xiaoran, et al.
Published: (2025)
TreeKV: Smooth Key-Value Cache Compression with Tree Structures
by: He, Ziwei, et al.
Published: (2025)
by: He, Ziwei, et al.
Published: (2025)
EchoKV: Efficient KV Cache Compression via Similarity-Based Reconstruction
by: Ji, Shiyu, et al.
Published: (2026)
by: Ji, Shiyu, et al.
Published: (2026)
Unlocking Data-free Low-bit Quantization with Matrix Decomposition for KV Cache Compression
by: Liu, Peiyu, et al.
Published: (2024)
by: Liu, Peiyu, et al.
Published: (2024)
Compressing Transformer Language Models via Matrix Product Operator Decomposition: A Case Study on PicoGPT
by: Javanmard, Younes, et al.
Published: (2026)
by: Javanmard, Younes, et al.
Published: (2026)
KV-Compress: Paged KV-Cache Compression with Variable Compression Rates per Attention Head
by: Rehg, Isaac
Published: (2024)
by: Rehg, Isaac
Published: (2024)
Which Heads Matter for Reasoning? RL-Guided KV Cache Compression
by: Du, Wenjie, et al.
Published: (2025)
by: Du, Wenjie, et al.
Published: (2025)
Understanding the Effects of Domain Finetuning on LLMs
by: Tanwar, Eshaan, et al.
Published: (2025)
by: Tanwar, Eshaan, et al.
Published: (2025)
Mitigating KV Cache Competition to Enhance User Experience in LLM Inference
by: Shen, Haiying, et al.
Published: (2025)
by: Shen, Haiying, et al.
Published: (2025)
Can LLMs replace Neil deGrasse Tyson? Evaluating the Reliability of LLMs as Science Communicators
by: Bajpai, Prasoon, et al.
Published: (2024)
by: Bajpai, Prasoon, et al.
Published: (2024)
KV-Distill: Nearly Lossless Learnable Context Compression for LLMs
by: Chari, Vivek, et al.
Published: (2025)
by: Chari, Vivek, et al.
Published: (2025)
Multilingual LLMs Inherently Reward In-Language Time-Sensitive Semantic Alignment for Low-Resource Languages
by: Bajpai, Ashutosh, et al.
Published: (2024)
by: Bajpai, Ashutosh, et al.
Published: (2024)
Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs
by: Ge, Suyu, et al.
Published: (2023)
by: Ge, Suyu, et al.
Published: (2023)
Harmonizing Code-mixed Conversations: Personality-assisted Code-mixed Response Generation in Dialogues
by: Kumar, Shivani, et al.
Published: (2024)
by: Kumar, Shivani, et al.
Published: (2024)
Decoding Memes: Benchmarking Narrative Role Classification across Multilingual and Multimodal Models
by: Sharma, Shivam, et al.
Published: (2025)
by: Sharma, Shivam, et al.
Published: (2025)
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs
by: Yang, Dongquan, et al.
Published: (2025)
by: Yang, Dongquan, et al.
Published: (2025)
Can LLMs reason over extended multilingual contexts? Towards long-context evaluation beyond retrieval and haystacks
by: Hengle, Amey, et al.
Published: (2025)
by: Hengle, Amey, et al.
Published: (2025)
Compactor: Calibrated Query-Agnostic KV Cache Compression with Approximate Leverage Scores
by: Chari, Vivek, et al.
Published: (2025)
by: Chari, Vivek, et al.
Published: (2025)
Latent Performance Profiling of Large Language Models
by: Chakraborty, Tanmoy, et al.
Published: (2026)
by: Chakraborty, Tanmoy, et al.
Published: (2026)
FreqKV: Key-Value Compression in Frequency Domain for Context Window Extension
by: Kai, Jushi, et al.
Published: (2025)
by: Kai, Jushi, et al.
Published: (2025)
Hate Personified: Investigating the role of LLMs in content moderation
by: Masud, Sarah, et al.
Published: (2024)
by: Masud, Sarah, et al.
Published: (2024)
Similar Items
-
Compression Laws for Large Language Models
by: Sengupta, Ayan, et al.
Published: (2025) -
You Only Prune Once: Designing Calibration-Free Model Compression With Policy Learning
by: Sengupta, Ayan, et al.
Published: (2025) -
Understanding the Physics of Key-Value Cache Compression for LLMs through Attention Dynamics
by: Ananthanarayanan, Samhruth, et al.
Published: (2026) -
Position: Enough of Scaling LLMs! Lets Focus on Downscaling
by: Goel, Yash, et al.
Published: (2025) -
The Art of Scaling Test-Time Compute for Large Language Models
by: Agarwal, Aradhye, et al.
Published: (2025)