KVCompose: Efficient Structured KV Cache Compression with Composite Tokens
Fuente:
arXiv
Saved in:
| Main Authors: | Akulov, Dmitry, Sana, Mohamed, De Domenico, Antonio, Salem, Tareq Si, Piovesan, Nicola, Ayed, Fadhel |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Goal-Oriented Time-Series Forecasting: Foundation Framework Design
by: Fechete, Luca-Andrei, et al.
Published: (2025)
by: Fechete, Luca-Andrei, et al.
Published: (2025)
HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime
by: Sana, Mohamed, et al.
Published: (2026)
by: Sana, Mohamed, et al.
Published: (2026)
Telecom Language Models: Must They Be Large?
by: Piovesan, Nicola, et al.
Published: (2024)
by: Piovesan, Nicola, et al.
Published: (2024)
TeleTables: A Benchmark for Large Language Models in Telecom Table Interpretation
by: Ezzakri, Anas, et al.
Published: (2025)
by: Ezzakri, Anas, et al.
Published: (2025)
Large Language Models for Telecom: Forthcoming Impact on the Industry
by: Maatouk, Ali, et al.
Published: (2023)
by: Maatouk, Ali, et al.
Published: (2023)
Telco-oRAG: Optimizing Retrieval-augmented Generation for Telecom Queries via Hybrid Retrieval and Neural Routing
by: Bornea, Andrei-Laurentiu, et al.
Published: (2025)
by: Bornea, Andrei-Laurentiu, et al.
Published: (2025)
TeleMath: A Benchmark for Large Language Models in Telecom Mathematical Problem Solving
by: Colle, Vincenzo, et al.
Published: (2025)
by: Colle, Vincenzo, et al.
Published: (2025)
Bandits in Flux: Adversarial Constraints in Dynamic Environments
by: Salem, Tareq Si
Published: (2026)
by: Salem, Tareq Si
Published: (2026)
Reasoning Language Models for Root Cause Analysis in 5G Wireless Networks
by: Sana, Mohamed, et al.
Published: (2025)
by: Sana, Mohamed, et al.
Published: (2025)
Beyond Token Eviction: Mixed-Dimension Budget Allocation for Efficient KV Cache Compression
by: Miao, Ruijie, et al.
Published: (2026)
by: Miao, Ruijie, et al.
Published: (2026)
ThinKV: Thought-Adaptive KV Cache Compression for Efficient Reasoning Models
by: Ramachandran, Akshat, et al.
Published: (2025)
by: Ramachandran, Akshat, et al.
Published: (2025)
Beyond the Chinese Restaurant and Pitman-Yor processes: Statistical Models with Double Power-law Behavior
by: Ayed, Fadhel, et al.
Published: (2019)
by: Ayed, Fadhel, et al.
Published: (2019)
ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification
by: He, Yefei, et al.
Published: (2024)
by: He, Yefei, et al.
Published: (2024)
The Pitfalls of KV Cache Compression
by: Chen, Alex, et al.
Published: (2025)
by: Chen, Alex, et al.
Published: (2025)
Training Transformers for KV Cache Compressibility
by: Gelberg, Yoav, et al.
Published: (2026)
by: Gelberg, Yoav, et al.
Published: (2026)
Telco-RAG: Navigating the Challenges of Retrieval-Augmented Language Models for Telecommunications
by: Bornea, Andrei-Laurentiu, et al.
Published: (2024)
by: Bornea, Andrei-Laurentiu, et al.
Published: (2024)
LouisKV: Efficient KV Cache Retrieval for Long Input-Output Sequences
by: Wu, Wenbo, et al.
Published: (2025)
by: Wu, Wenbo, et al.
Published: (2025)
LongFlow: Efficient KV Cache Compression for Reasoning Models
by: Su, Yi, et al.
Published: (2026)
by: Su, Yi, et al.
Published: (2026)
Palu: Compressing KV-Cache with Low-Rank Projection
by: Chang, Chi-Chih, et al.
Published: (2024)
by: Chang, Chi-Chih, et al.
Published: (2024)
Hurwitz Quaternion Multiplicative Quantization for KV Cache Compression
by: Swain, Kabir, et al.
Published: (2026)
by: Swain, Kabir, et al.
Published: (2026)
SmallKV: Small Model Assisted Compensation of KV Cache Compression for Efficient LLM Inference
by: Zhao, Yi, et al.
Published: (2025)
by: Zhao, Yi, et al.
Published: (2025)
KV-CAR: KV Cache Compression using Autoencoders and KV Reuse in Large Language Models
by: Roy, Sourjya, et al.
Published: (2025)
by: Roy, Sourjya, et al.
Published: (2025)
Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression
by: Li, Kunjun, et al.
Published: (2025)
by: Li, Kunjun, et al.
Published: (2025)
KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
by: Tian, Yuxuan, et al.
Published: (2025)
by: Tian, Yuxuan, et al.
Published: (2025)
RazorAttention: Efficient KV Cache Compression Through Retrieval Heads
by: Tang, Hanlin, et al.
Published: (2024)
by: Tang, Hanlin, et al.
Published: (2024)
Data-driven Energy Efficiency Modelling in Large-scale Networks: An Expert Knowledge and ML-based Approach
by: López-Pérez, David, et al.
Published: (2023)
by: López-Pérez, David, et al.
Published: (2023)
KVSculpt: KV Cache Compression as Distillation
by: Jiang, Bo, et al.
Published: (2026)
by: Jiang, Bo, et al.
Published: (2026)
CommonKV: Compressing KV Cache with Cross-layer Parameter Sharing
by: Wang, Yixuan, et al.
Published: (2025)
by: Wang, Yixuan, et al.
Published: (2025)
One Size Does Not Fit All: Token-Wise Adaptive Compression for KV Cache
by: Lu, Liming, et al.
Published: (2026)
by: Lu, Liming, et al.
Published: (2026)
xKV: Cross-Layer KV-Cache Compression via Aligned Singular Vector Extraction
by: Chang, Chi-Chih, et al.
Published: (2025)
by: Chang, Chi-Chih, et al.
Published: (2025)
EvolKV: Evolutionary KV Cache Compression for LLM Inference
by: Yu, Bohan, et al.
Published: (2025)
by: Yu, Bohan, et al.
Published: (2025)
Cache What Lasts: Token Retention for Memory-Bounded KV Cache in LLMs
by: Bui, Ngoc, et al.
Published: (2025)
by: Bui, Ngoc, et al.
Published: (2025)
No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization
by: Yang, June Yong, et al.
Published: (2024)
by: Yang, June Yong, et al.
Published: (2024)
EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
by: Feng, Shaoting, et al.
Published: (2025)
by: Feng, Shaoting, et al.
Published: (2025)
KV Pareto: Systems-Level Optimization of KV Cache and Model Compression for Long Context Inference
by: Gokhale, Sai, et al.
Published: (2025)
by: Gokhale, Sai, et al.
Published: (2025)
Inference-Time Hyper-Scaling with KV Cache Compression
by: Łańcucki, Adrian, et al.
Published: (2025)
by: Łańcucki, Adrian, et al.
Published: (2025)
ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression
by: Liu, Guangda, et al.
Published: (2024)
by: Liu, Guangda, et al.
Published: (2024)
KVzip: Query-Agnostic KV Cache Compression with Context Reconstruction
by: Kim, Jang-Hyun, et al.
Published: (2025)
by: Kim, Jang-Hyun, et al.
Published: (2025)
FastKV: Decoupling of Context Reduction and KV Cache Compression for Prefill-Decoding Acceleration
by: Jo, Dongwon, et al.
Published: (2025)
by: Jo, Dongwon, et al.
Published: (2025)
Get More with LESS: Synthesizing Recurrence with KV Cache Compression for Efficient LLM Inference
by: Dong, Harry, et al.
Published: (2024)
by: Dong, Harry, et al.
Published: (2024)
Similar Items
-
Goal-Oriented Time-Series Forecasting: Foundation Framework Design
by: Fechete, Luca-Andrei, et al.
Published: (2025) -
HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime
by: Sana, Mohamed, et al.
Published: (2026) -
Telecom Language Models: Must They Be Large?
by: Piovesan, Nicola, et al.
Published: (2024) -
TeleTables: A Benchmark for Large Language Models in Telecom Table Interpretation
by: Ezzakri, Anas, et al.
Published: (2025) -
Large Language Models for Telecom: Forthcoming Impact on the Industry
by: Maatouk, Ali, et al.
Published: (2023)