KVzip: Query-Agnostic KV Cache Compression with Context Reconstruction
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Jang-Hyun, Kim, Jinuk, Kwon, Sangwoo, Lee, Jae W., Yun, Sangdoo, Song, Hyun Oh |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Fast KVzip: Efficient and Accurate LLM Inference with Gated KV Eviction
by: Kim, Jang-Hyun, et al.
Published: (2026)
by: Kim, Jang-Hyun, et al.
Published: (2026)
Compressed Context Memory For Online Language Model Interaction
by: Kim, Jang-Hyun, et al.
Published: (2023)
by: Kim, Jang-Hyun, et al.
Published: (2023)
FastKV: Decoupling of Context Reduction and KV Cache Compression for Prefill-Decoding Acceleration
by: Jo, Dongwon, et al.
Published: (2025)
by: Jo, Dongwon, et al.
Published: (2025)
ParisKV: Fast and Drift-Robust KV-Cache Retrieval for Long-Context LLMs
by: Qi, Yanlin, et al.
Published: (2026)
by: Qi, Yanlin, et al.
Published: (2026)
Large-Scale Targeted Cause Discovery via Learning from Simulated Data
by: Kim, Jang-Hyun, et al.
Published: (2024)
by: Kim, Jang-Hyun, et al.
Published: (2024)
Rethinking Caching for LLM Serving Systems: Beyond Traditional Heuristics
by: Kim, Jungwoo, et al.
Published: (2025)
by: Kim, Jungwoo, et al.
Published: (2025)
LayerMerge: Neural Network Depth Compression through Layer Pruning and Merging
by: Kim, Jinuk, et al.
Published: (2024)
by: Kim, Jinuk, et al.
Published: (2024)
Hydro: Adaptive Query Processing of ML Queries
by: Kakkar, Gaurav Tarlok, et al.
Published: (2024)
by: Kakkar, Gaurav Tarlok, et al.
Published: (2024)
Effective Dataset Distillation for Spatio-Temporal Forecasting with Bi-dimensional Compression
by: Kwon, Taehyung, et al.
Published: (2026)
by: Kwon, Taehyung, et al.
Published: (2026)
RELOAD: A Robust and Efficient Learned Query Optimizer for Database Systems
by: Lee, Seokwon, et al.
Published: (2026)
by: Lee, Seokwon, et al.
Published: (2026)
Rule2DRC: Benchmarking LLM Agents for DRC Script Synthesis with Execution-Guided Test Generation
by: Kim, Jinuk, et al.
Published: (2026)
by: Kim, Jinuk, et al.
Published: (2026)
GuidedQuant: Large Language Model Quantization via Exploiting End Loss Guidance
by: Kim, Jinuk, et al.
Published: (2025)
by: Kim, Jinuk, et al.
Published: (2025)
FibQuant: Universal Vector Quantization for Random-Access KV-Cache Compression
by: Lee, Namyoon, et al.
Published: (2026)
by: Lee, Namyoon, et al.
Published: (2026)
No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization
by: Yang, June Yong, et al.
Published: (2024)
by: Yang, June Yong, et al.
Published: (2024)
Task-Agnostic Contrastive Pretraining for Relational Deep Learning
by: Peleška, Jakub, et al.
Published: (2025)
by: Peleška, Jakub, et al.
Published: (2025)
Data-Agnostic Cardinality Learning from Imperfect Workloads
by: Wu, Peizhi, et al.
Published: (2025)
by: Wu, Peizhi, et al.
Published: (2025)
Private Queries with Sigma-Counting
by: Gao, Jun, et al.
Published: (2025)
by: Gao, Jun, et al.
Published: (2025)
Polynomial Context-Truncation Sensitivity in Autoregressive Language Models: Sequential Wyner-Ziv Bounds for KV Cache Compression
by: Kim, Munsik
Published: (2026)
by: Kim, Munsik
Published: (2026)
Multi-View Graph Convolution Network for Internal Talent Recommendation Based on Enterprise Emails
by: Kim, Soo Hyun, et al.
Published: (2025)
by: Kim, Soo Hyun, et al.
Published: (2025)
The Unreasonable Effectiveness of LLMs for Query Optimization
by: Akioyamen, Peter, et al.
Published: (2024)
by: Akioyamen, Peter, et al.
Published: (2024)
LRAgent: Efficient KV Cache Sharing for Multi-LoRA LLM Agents
by: Jeon, Hyesung, et al.
Published: (2026)
by: Jeon, Hyesung, et al.
Published: (2026)
Low Rank Learning for Offline Query Optimization
by: Yi, Zixuan, et al.
Published: (2025)
by: Yi, Zixuan, et al.
Published: (2025)
LookaheadKV: Fast and Accurate KV Cache Eviction by Glimpsing into the Future without Generation
by: Ahn, Jinwoo, et al.
Published: (2026)
by: Ahn, Jinwoo, et al.
Published: (2026)
Predictive Query-based Pipeline for Graph Data
by: Neto, Plácido A Souza
Published: (2024)
by: Neto, Plácido A Souza
Published: (2024)
Incorporating Deep Learning Design in Database Queries
by: Lubarsky, Yuval Lev, et al.
Published: (2026)
by: Lubarsky, Yuval Lev, et al.
Published: (2026)
Adversarial Query Synthesis via Bayesian Optimization
by: Tao, Jeffrey, et al.
Published: (2026)
by: Tao, Jeffrey, et al.
Published: (2026)
Sibyl: Forecasting Time-Evolving Query Workloads
by: Huang, Hanxian, et al.
Published: (2024)
by: Huang, Hanxian, et al.
Published: (2024)
KV Pareto: Systems-Level Optimization of KV Cache and Model Compression for Long Context Inference
by: Gokhale, Sai, et al.
Published: (2025)
by: Gokhale, Sai, et al.
Published: (2025)
ACE: A Cardinality Estimator for Set-Valued Queries
by: Sheng, Yufan, et al.
Published: (2025)
by: Sheng, Yufan, et al.
Published: (2025)
Optimizing LLM Queries in Relational Data Analytics Workloads
by: Liu, Shu, et al.
Published: (2024)
by: Liu, Shu, et al.
Published: (2024)
A Declarative Query Language for Scientific Machine Learning
by: Jamil, Hasan M
Published: (2024)
by: Jamil, Hasan M
Published: (2024)
Leveraging Approximate Caching for Faster Retrieval-Augmented Generation
by: Bergman, Shai, et al.
Published: (2025)
by: Bergman, Shai, et al.
Published: (2025)
SemBench: A Benchmark for Semantic Query Processing Engines
by: Lao, Jiale, et al.
Published: (2025)
by: Lao, Jiale, et al.
Published: (2025)
OjaKV: Context-Aware Online Low-Rank KV Cache Compression
by: Zhu, Yuxuan, et al.
Published: (2025)
by: Zhu, Yuxuan, et al.
Published: (2025)
Accelerating Storage-Based Training for Graph Neural Networks
by: Jang, Myung-Hwan, et al.
Published: (2026)
by: Jang, Myung-Hwan, et al.
Published: (2026)
Training-Free Query Optimization via LLM-Based Plan Similarity
by: Vasilenko, Nikita, et al.
Published: (2025)
by: Vasilenko, Nikita, et al.
Published: (2025)
Compass: SLO-aware Query Planner for Compound AI Serving at Scale
by: Liu, Banruo, et al.
Published: (2025)
by: Liu, Banruo, et al.
Published: (2025)
LearnedWMP: Workload Memory Prediction Using Distribution of Query Templates
by: Quader, Shaikh, et al.
Published: (2024)
by: Quader, Shaikh, et al.
Published: (2024)
Adaptive KV-Cache Compression without Manually Setting Budget
by: Tang, Chenxia, et al.
Published: (2025)
by: Tang, Chenxia, et al.
Published: (2025)
The Pitfalls of KV Cache Compression
by: Chen, Alex, et al.
Published: (2025)
by: Chen, Alex, et al.
Published: (2025)
Similar Items
-
Fast KVzip: Efficient and Accurate LLM Inference with Gated KV Eviction
by: Kim, Jang-Hyun, et al.
Published: (2026) -
Compressed Context Memory For Online Language Model Interaction
by: Kim, Jang-Hyun, et al.
Published: (2023) -
FastKV: Decoupling of Context Reduction and KV Cache Compression for Prefill-Decoding Acceleration
by: Jo, Dongwon, et al.
Published: (2025) -
ParisKV: Fast and Drift-Robust KV-Cache Retrieval for Long-Context LLMs
by: Qi, Yanlin, et al.
Published: (2026) -
Large-Scale Targeted Cause Discovery via Learning from Simulated Data
by: Kim, Jang-Hyun, et al.
Published: (2024)