Adaptive KV-Cache Compression without Manually Setting Budget
Fuente:
arXiv
Salvato in:
| Autori principali: | Tang, Chenxia, Liu, Jianchun, Xu, Hongli, Huang, Liusheng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Accelerating Mixture-of-Expert Inference with Adaptive Expert Split Mechanism
di: Yan, Jiaming, et al.
Pubblicazione: (2025)
di: Yan, Jiaming, et al.
Pubblicazione: (2025)
Improving LLM Reasoning via Dependency-Aware Query Decomposition and Logic-Parallel Content Expansion
di: Gao, Xianjun, et al.
Pubblicazione: (2025)
di: Gao, Xianjun, et al.
Pubblicazione: (2025)
Top-$nσ$: Not All Logits Are You Need
di: Tang, Chenxia, et al.
Pubblicazione: (2024)
di: Tang, Chenxia, et al.
Pubblicazione: (2024)
WAter: A Workload-Adaptive Knob Tuning System based on Workload Compression
di: Wang, Yibo, et al.
Pubblicazione: (2026)
di: Wang, Yibo, et al.
Pubblicazione: (2026)
EHL*: Memory-Budgeted Indexing for Ultrafast Optimal Euclidean Pathfinding
di: Du, Jinchun, et al.
Pubblicazione: (2024)
di: Du, Jinchun, et al.
Pubblicazione: (2024)
QCFuse: Query-Centric Cache Fusion for Efficient RAG Inference
di: Yan, Jianxin, et al.
Pubblicazione: (2026)
di: Yan, Jianxin, et al.
Pubblicazione: (2026)
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
di: Feng, Yuan, et al.
Pubblicazione: (2024)
di: Feng, Yuan, et al.
Pubblicazione: (2024)
SABlock: Semantic-Aware KV Cache Eviction with Adaptive Compression Block Size
di: Chen, Jinhan, et al.
Pubblicazione: (2025)
di: Chen, Jinhan, et al.
Pubblicazione: (2025)
Category-Aware Semantic Caching for Heterogeneous LLM Workloads
di: Wang, Chen, et al.
Pubblicazione: (2025)
di: Wang, Chen, et al.
Pubblicazione: (2025)
Towards Communication-Efficient Decentralized Federated Graph Learning over Non-IID Data
di: Wang, Shilong, et al.
Pubblicazione: (2025)
di: Wang, Shilong, et al.
Pubblicazione: (2025)
The Pitfalls of KV Cache Compression
di: Chen, Alex, et al.
Pubblicazione: (2025)
di: Chen, Alex, et al.
Pubblicazione: (2025)
MontePrep: Monte-Carlo-Driven Automatic Data Preparation without Target Data Instances
di: Ge, Congcong, et al.
Pubblicazione: (2025)
di: Ge, Congcong, et al.
Pubblicazione: (2025)
Budget-aware Query Tuning: An AutoML Perspective
di: Wu, Wentao, et al.
Pubblicazione: (2024)
di: Wu, Wentao, et al.
Pubblicazione: (2024)
KVzip: Query-Agnostic KV Cache Compression with Context Reconstruction
di: Kim, Jang-Hyun, et al.
Pubblicazione: (2025)
di: Kim, Jang-Hyun, et al.
Pubblicazione: (2025)
Lossless KV Cache Compression to 2%
di: Yang, Zhen, et al.
Pubblicazione: (2024)
di: Yang, Zhen, et al.
Pubblicazione: (2024)
Enhancing Large Multimodal Models with Adaptive Sparsity and KV Cache Compression
di: Zhang, Te, et al.
Pubblicazione: (2025)
di: Zhang, Te, et al.
Pubblicazione: (2025)
DumpKV: Learning based lifetime aware garbage collection for key value separation in LSM-tree
di: Zhuang, Zhutao, et al.
Pubblicazione: (2024)
di: Zhuang, Zhutao, et al.
Pubblicazione: (2024)
Reliable Text-to-SQL with Adaptive Abstention
di: Chen, Kaiwen, et al.
Pubblicazione: (2025)
di: Chen, Kaiwen, et al.
Pubblicazione: (2025)
PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
di: Cai, Zefan, et al.
Pubblicazione: (2024)
di: Cai, Zefan, et al.
Pubblicazione: (2024)
Towards Automated Cross-domain Exploratory Data Analysis through Large Language Models
di: Zhu, Jun-Peng, et al.
Pubblicazione: (2024)
di: Zhu, Jun-Peng, et al.
Pubblicazione: (2024)
ZSMerge: Zero-Shot KV Cache Compression for Memory-Efficient Long-Context LLMs
di: Liu, Xin, et al.
Pubblicazione: (2025)
di: Liu, Xin, et al.
Pubblicazione: (2025)
Graph Pattern-based Association Rules Evaluated Under No-repeated-anything Semantics in the Graph Transactional Setting
di: Ell, Basil
Pubblicazione: (2025)
di: Ell, Basil
Pubblicazione: (2025)
DeltaKV: Residual-Based KV Cache Compression via Long-Range Similarity
di: Hao, Jitai, et al.
Pubblicazione: (2026)
di: Hao, Jitai, et al.
Pubblicazione: (2026)
On 10x Better Scalability: KV Stores Scale Up KV Cache
di: Yu, Weiping, et al.
Pubblicazione: (2025)
di: Yu, Weiping, et al.
Pubblicazione: (2025)
Meta-Soft: Leveraging Composable Meta-Tokens for Context-Preserving KV Cache Compression
di: Luo, Wei, et al.
Pubblicazione: (2026)
di: Luo, Wei, et al.
Pubblicazione: (2026)
Cardinality Estimation for High Dimensional Similarity Queries with Adaptive Bucket Probing
di: Chen, Zhonghan, et al.
Pubblicazione: (2026)
di: Chen, Zhonghan, et al.
Pubblicazione: (2026)
CommonKV: Compressing KV Cache with Cross-layer Parameter Sharing
di: Wang, Yixuan, et al.
Pubblicazione: (2025)
di: Wang, Yixuan, et al.
Pubblicazione: (2025)
R-KV: Redundancy-aware KV Cache Compression for Reasoning Models
di: Cai, Zefan, et al.
Pubblicazione: (2025)
di: Cai, Zefan, et al.
Pubblicazione: (2025)
KVSculpt: KV Cache Compression as Distillation
di: Jiang, Bo, et al.
Pubblicazione: (2026)
di: Jiang, Bo, et al.
Pubblicazione: (2026)
Aixel: A Unified, Adaptive and Extensible System for AI-powered Data Analysis
di: Zhang, Meihui, et al.
Pubblicazione: (2025)
di: Zhang, Meihui, et al.
Pubblicazione: (2025)
Palu: Compressing KV-Cache with Low-Rank Projection
di: Chang, Chi-Chih, et al.
Pubblicazione: (2024)
di: Chang, Chi-Chih, et al.
Pubblicazione: (2024)
ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression
di: Liu, Guangda, et al.
Pubblicazione: (2024)
di: Liu, Guangda, et al.
Pubblicazione: (2024)
Moment-KV: Momentum-Based Decode-Time KV Cache Compression for Long Generation
di: Jana, Soumyadeep, et al.
Pubblicazione: (2026)
di: Jana, Soumyadeep, et al.
Pubblicazione: (2026)
FlexiDataGen: An Adaptive LLM Framework for Dynamic Semantic Dataset Generation in Sensitive Domains
di: Jelodar, Hamed, et al.
Pubblicazione: (2025)
di: Jelodar, Hamed, et al.
Pubblicazione: (2025)
DistJoin: A Decoupled Join Cardinality Estimator based on Adaptive Neural Predicate Modulation
di: Zhang, Kaixin, et al.
Pubblicazione: (2025)
di: Zhang, Kaixin, et al.
Pubblicazione: (2025)
STaR-KV: Spatio-Temporal Adaptive Re-weighting for KV Cache Compression in GUI Vision-Language Models
di: Han, Yuhang, et al.
Pubblicazione: (2026)
di: Han, Yuhang, et al.
Pubblicazione: (2026)
FedQuad: Adaptive Layer-wise LoRA Deployment and Activation Quantization for Federated Fine-Tuning
di: Li, Rukuo, et al.
Pubblicazione: (2025)
di: Li, Rukuo, et al.
Pubblicazione: (2025)
LAVa: Layer-wise KV Cache Eviction with Dynamic Budget Allocation
di: Shen, Yiqun, et al.
Pubblicazione: (2025)
di: Shen, Yiqun, et al.
Pubblicazione: (2025)
MiniCache: KV Cache Compression in Depth Dimension for Large Language Models
di: Liu, Akide, et al.
Pubblicazione: (2024)
di: Liu, Akide, et al.
Pubblicazione: (2024)
GRACE: A Dynamic Coreset Selection Framework for Large Language Model Optimization
di: Tang, Tianhao, et al.
Pubblicazione: (2026)
di: Tang, Tianhao, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Accelerating Mixture-of-Expert Inference with Adaptive Expert Split Mechanism
di: Yan, Jiaming, et al.
Pubblicazione: (2025) -
Improving LLM Reasoning via Dependency-Aware Query Decomposition and Logic-Parallel Content Expansion
di: Gao, Xianjun, et al.
Pubblicazione: (2025) -
Top-$nσ$: Not All Logits Are You Need
di: Tang, Chenxia, et al.
Pubblicazione: (2024) -
WAter: A Workload-Adaptive Knob Tuning System based on Workload Compression
di: Wang, Yibo, et al.
Pubblicazione: (2026) -
EHL*: Memory-Budgeted Indexing for Ultrafast Optimal Euclidean Pathfinding
di: Du, Jinchun, et al.
Pubblicazione: (2024)