Adaptive KV-Cache Compression without Manually Setting Budget
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tang, Chenxia, Liu, Jianchun, Xu, Hongli, Huang, Liusheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Accelerating Mixture-of-Expert Inference with Adaptive Expert Split Mechanism
von: Yan, Jiaming, et al.
Veröffentlicht: (2025)
von: Yan, Jiaming, et al.
Veröffentlicht: (2025)
Improving LLM Reasoning via Dependency-Aware Query Decomposition and Logic-Parallel Content Expansion
von: Gao, Xianjun, et al.
Veröffentlicht: (2025)
von: Gao, Xianjun, et al.
Veröffentlicht: (2025)
Top-$nσ$: Not All Logits Are You Need
von: Tang, Chenxia, et al.
Veröffentlicht: (2024)
von: Tang, Chenxia, et al.
Veröffentlicht: (2024)
WAter: A Workload-Adaptive Knob Tuning System based on Workload Compression
von: Wang, Yibo, et al.
Veröffentlicht: (2026)
von: Wang, Yibo, et al.
Veröffentlicht: (2026)
EHL*: Memory-Budgeted Indexing for Ultrafast Optimal Euclidean Pathfinding
von: Du, Jinchun, et al.
Veröffentlicht: (2024)
von: Du, Jinchun, et al.
Veröffentlicht: (2024)
QCFuse: Query-Centric Cache Fusion for Efficient RAG Inference
von: Yan, Jianxin, et al.
Veröffentlicht: (2026)
von: Yan, Jianxin, et al.
Veröffentlicht: (2026)
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
von: Feng, Yuan, et al.
Veröffentlicht: (2024)
von: Feng, Yuan, et al.
Veröffentlicht: (2024)
SABlock: Semantic-Aware KV Cache Eviction with Adaptive Compression Block Size
von: Chen, Jinhan, et al.
Veröffentlicht: (2025)
von: Chen, Jinhan, et al.
Veröffentlicht: (2025)
Category-Aware Semantic Caching for Heterogeneous LLM Workloads
von: Wang, Chen, et al.
Veröffentlicht: (2025)
von: Wang, Chen, et al.
Veröffentlicht: (2025)
Towards Communication-Efficient Decentralized Federated Graph Learning over Non-IID Data
von: Wang, Shilong, et al.
Veröffentlicht: (2025)
von: Wang, Shilong, et al.
Veröffentlicht: (2025)
The Pitfalls of KV Cache Compression
von: Chen, Alex, et al.
Veröffentlicht: (2025)
von: Chen, Alex, et al.
Veröffentlicht: (2025)
MontePrep: Monte-Carlo-Driven Automatic Data Preparation without Target Data Instances
von: Ge, Congcong, et al.
Veröffentlicht: (2025)
von: Ge, Congcong, et al.
Veröffentlicht: (2025)
Budget-aware Query Tuning: An AutoML Perspective
von: Wu, Wentao, et al.
Veröffentlicht: (2024)
von: Wu, Wentao, et al.
Veröffentlicht: (2024)
KVzip: Query-Agnostic KV Cache Compression with Context Reconstruction
von: Kim, Jang-Hyun, et al.
Veröffentlicht: (2025)
von: Kim, Jang-Hyun, et al.
Veröffentlicht: (2025)
Lossless KV Cache Compression to 2%
von: Yang, Zhen, et al.
Veröffentlicht: (2024)
von: Yang, Zhen, et al.
Veröffentlicht: (2024)
Enhancing Large Multimodal Models with Adaptive Sparsity and KV Cache Compression
von: Zhang, Te, et al.
Veröffentlicht: (2025)
von: Zhang, Te, et al.
Veröffentlicht: (2025)
DumpKV: Learning based lifetime aware garbage collection for key value separation in LSM-tree
von: Zhuang, Zhutao, et al.
Veröffentlicht: (2024)
von: Zhuang, Zhutao, et al.
Veröffentlicht: (2024)
Reliable Text-to-SQL with Adaptive Abstention
von: Chen, Kaiwen, et al.
Veröffentlicht: (2025)
von: Chen, Kaiwen, et al.
Veröffentlicht: (2025)
PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
von: Cai, Zefan, et al.
Veröffentlicht: (2024)
von: Cai, Zefan, et al.
Veröffentlicht: (2024)
Towards Automated Cross-domain Exploratory Data Analysis through Large Language Models
von: Zhu, Jun-Peng, et al.
Veröffentlicht: (2024)
von: Zhu, Jun-Peng, et al.
Veröffentlicht: (2024)
ZSMerge: Zero-Shot KV Cache Compression for Memory-Efficient Long-Context LLMs
von: Liu, Xin, et al.
Veröffentlicht: (2025)
von: Liu, Xin, et al.
Veröffentlicht: (2025)
Graph Pattern-based Association Rules Evaluated Under No-repeated-anything Semantics in the Graph Transactional Setting
von: Ell, Basil
Veröffentlicht: (2025)
von: Ell, Basil
Veröffentlicht: (2025)
DeltaKV: Residual-Based KV Cache Compression via Long-Range Similarity
von: Hao, Jitai, et al.
Veröffentlicht: (2026)
von: Hao, Jitai, et al.
Veröffentlicht: (2026)
On 10x Better Scalability: KV Stores Scale Up KV Cache
von: Yu, Weiping, et al.
Veröffentlicht: (2025)
von: Yu, Weiping, et al.
Veröffentlicht: (2025)
Meta-Soft: Leveraging Composable Meta-Tokens for Context-Preserving KV Cache Compression
von: Luo, Wei, et al.
Veröffentlicht: (2026)
von: Luo, Wei, et al.
Veröffentlicht: (2026)
Cardinality Estimation for High Dimensional Similarity Queries with Adaptive Bucket Probing
von: Chen, Zhonghan, et al.
Veröffentlicht: (2026)
von: Chen, Zhonghan, et al.
Veröffentlicht: (2026)
CommonKV: Compressing KV Cache with Cross-layer Parameter Sharing
von: Wang, Yixuan, et al.
Veröffentlicht: (2025)
von: Wang, Yixuan, et al.
Veröffentlicht: (2025)
R-KV: Redundancy-aware KV Cache Compression for Reasoning Models
von: Cai, Zefan, et al.
Veröffentlicht: (2025)
von: Cai, Zefan, et al.
Veröffentlicht: (2025)
KVSculpt: KV Cache Compression as Distillation
von: Jiang, Bo, et al.
Veröffentlicht: (2026)
von: Jiang, Bo, et al.
Veröffentlicht: (2026)
Aixel: A Unified, Adaptive and Extensible System for AI-powered Data Analysis
von: Zhang, Meihui, et al.
Veröffentlicht: (2025)
von: Zhang, Meihui, et al.
Veröffentlicht: (2025)
Palu: Compressing KV-Cache with Low-Rank Projection
von: Chang, Chi-Chih, et al.
Veröffentlicht: (2024)
von: Chang, Chi-Chih, et al.
Veröffentlicht: (2024)
ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression
von: Liu, Guangda, et al.
Veröffentlicht: (2024)
von: Liu, Guangda, et al.
Veröffentlicht: (2024)
Moment-KV: Momentum-Based Decode-Time KV Cache Compression for Long Generation
von: Jana, Soumyadeep, et al.
Veröffentlicht: (2026)
von: Jana, Soumyadeep, et al.
Veröffentlicht: (2026)
FedQuad: Adaptive Layer-wise LoRA Deployment and Activation Quantization for Federated Fine-Tuning
von: Li, Rukuo, et al.
Veröffentlicht: (2025)
von: Li, Rukuo, et al.
Veröffentlicht: (2025)
FlexiDataGen: An Adaptive LLM Framework for Dynamic Semantic Dataset Generation in Sensitive Domains
von: Jelodar, Hamed, et al.
Veröffentlicht: (2025)
von: Jelodar, Hamed, et al.
Veröffentlicht: (2025)
DistJoin: A Decoupled Join Cardinality Estimator based on Adaptive Neural Predicate Modulation
von: Zhang, Kaixin, et al.
Veröffentlicht: (2025)
von: Zhang, Kaixin, et al.
Veröffentlicht: (2025)
STaR-KV: Spatio-Temporal Adaptive Re-weighting for KV Cache Compression in GUI Vision-Language Models
von: Han, Yuhang, et al.
Veröffentlicht: (2026)
von: Han, Yuhang, et al.
Veröffentlicht: (2026)
LAVa: Layer-wise KV Cache Eviction with Dynamic Budget Allocation
von: Shen, Yiqun, et al.
Veröffentlicht: (2025)
von: Shen, Yiqun, et al.
Veröffentlicht: (2025)
MiniCache: KV Cache Compression in Depth Dimension for Large Language Models
von: Liu, Akide, et al.
Veröffentlicht: (2024)
von: Liu, Akide, et al.
Veröffentlicht: (2024)
GRACE: A Dynamic Coreset Selection Framework for Large Language Model Optimization
von: Tang, Tianhao, et al.
Veröffentlicht: (2026)
von: Tang, Tianhao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Accelerating Mixture-of-Expert Inference with Adaptive Expert Split Mechanism
von: Yan, Jiaming, et al.
Veröffentlicht: (2025) -
Improving LLM Reasoning via Dependency-Aware Query Decomposition and Logic-Parallel Content Expansion
von: Gao, Xianjun, et al.
Veröffentlicht: (2025) -
Top-$nσ$: Not All Logits Are You Need
von: Tang, Chenxia, et al.
Veröffentlicht: (2024) -
WAter: A Workload-Adaptive Knob Tuning System based on Workload Compression
von: Wang, Yibo, et al.
Veröffentlicht: (2026) -
EHL*: Memory-Budgeted Indexing for Ultrafast Optimal Euclidean Pathfinding
von: Du, Jinchun, et al.
Veröffentlicht: (2024)