KV-Distill: Nearly Lossless Learnable Context Compression for LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chari, Vivek, Qin, Guanghui, Van Durme, Benjamin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Compactor: Calibrated Query-Agnostic KV Cache Compression with Approximate Leverage Scores
von: Chari, Vivek, et al.
Veröffentlicht: (2025)
von: Chari, Vivek, et al.
Veröffentlicht: (2025)
Lossless KV Cache Compression to 2%
von: Yang, Zhen, et al.
Veröffentlicht: (2024)
von: Yang, Zhen, et al.
Veröffentlicht: (2024)
GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
von: Kang, Hao, et al.
Veröffentlicht: (2024)
von: Kang, Hao, et al.
Veröffentlicht: (2024)
KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
von: Tian, Yuxuan, et al.
Veröffentlicht: (2025)
von: Tian, Yuxuan, et al.
Veröffentlicht: (2025)
Dodo: Dynamic Contextual Compression for Decoder-only LMs
von: Qin, Guanghui, et al.
Veröffentlicht: (2023)
von: Qin, Guanghui, et al.
Veröffentlicht: (2023)
ZSMerge: Zero-Shot KV Cache Compression for Memory-Efficient Long-Context LLMs
von: Liu, Xin, et al.
Veröffentlicht: (2025)
von: Liu, Xin, et al.
Veröffentlicht: (2025)
KVSculpt: KV Cache Compression as Distillation
von: Jiang, Bo, et al.
Veröffentlicht: (2026)
von: Jiang, Bo, et al.
Veröffentlicht: (2026)
SEQR: Secure and Efficient QR-based LoRA Routing
von: Fleshman, William, et al.
Veröffentlicht: (2025)
von: Fleshman, William, et al.
Veröffentlicht: (2025)
LoRA-Augmented Generation (LAG) for Knowledge-Intensive Language Tasks
von: Fleshman, William, et al.
Veröffentlicht: (2025)
von: Fleshman, William, et al.
Veröffentlicht: (2025)
SpectR: Dynamically Composing LM Experts with Spectral Routing
von: Fleshman, William, et al.
Veröffentlicht: (2025)
von: Fleshman, William, et al.
Veröffentlicht: (2025)
RE-Adapt: Reverse Engineered Adaptation of Large Language Models
von: Fleshman, William, et al.
Veröffentlicht: (2024)
von: Fleshman, William, et al.
Veröffentlicht: (2024)
Interpreting User Requests in the Context of Natural Language Standing Instructions
von: Moghe, Nikita, et al.
Veröffentlicht: (2023)
von: Moghe, Nikita, et al.
Veröffentlicht: (2023)
LM Agents for Coordinating Multi-User Information Gathering
von: Jhamtani, Harsh, et al.
Veröffentlicht: (2025)
von: Jhamtani, Harsh, et al.
Veröffentlicht: (2025)
AdapterSwap: Continuous Training of LLMs with Data Removal and Access-Control Guarantees
von: Fleshman, William, et al.
Veröffentlicht: (2024)
von: Fleshman, William, et al.
Veröffentlicht: (2024)
OjaKV: Context-Aware Online Low-Rank KV Cache Compression
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
FreqKV: Key-Value Compression in Frequency Domain for Context Window Extension
von: Kai, Jushi, et al.
Veröffentlicht: (2025)
von: Kai, Jushi, et al.
Veröffentlicht: (2025)
Language Models and Logic Programs for Trustworthy Tax Reasoning
von: Jurayj, William, et al.
Veröffentlicht: (2025)
von: Jurayj, William, et al.
Veröffentlicht: (2025)
KVTuner: Sensitivity-Aware Layer-Wise Mixed-Precision KV Cache Quantization for Efficient and Nearly Lossless LLM Inference
von: Li, Xing, et al.
Veröffentlicht: (2025)
von: Li, Xing, et al.
Veröffentlicht: (2025)
LLMs in the Imaginarium: Tool Learning through Simulated Trial and Error
von: Wang, Boshi, et al.
Veröffentlicht: (2024)
von: Wang, Boshi, et al.
Veröffentlicht: (2024)
RE-AdaptIR: Improving Information Retrieval through Reverse Engineered Adaptation
von: Fleshman, William, et al.
Veröffentlicht: (2024)
von: Fleshman, William, et al.
Veröffentlicht: (2024)
R-KV: Redundancy-aware KV Cache Compression for Reasoning Models
von: Cai, Zefan, et al.
Veröffentlicht: (2025)
von: Cai, Zefan, et al.
Veröffentlicht: (2025)
PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
von: Cai, Zefan, et al.
Veröffentlicht: (2024)
von: Cai, Zefan, et al.
Veröffentlicht: (2024)
SELF-[IN]CORRECT: LLMs Struggle with Discriminating Self-Generated Responses
von: Jiang, Dongwei, et al.
Veröffentlicht: (2024)
von: Jiang, Dongwei, et al.
Veröffentlicht: (2024)
DeltaKV: Residual-Based KV Cache Compression via Long-Range Similarity
von: Hao, Jitai, et al.
Veröffentlicht: (2026)
von: Hao, Jitai, et al.
Veröffentlicht: (2026)
Crystal: Characterizing Relative Impact of Scholarly Publications
von: Collison, Hannah, et al.
Veröffentlicht: (2026)
von: Collison, Hannah, et al.
Veröffentlicht: (2026)
HeteroCache: A Dynamic Retrieval Approach to Heterogeneous KV Cache Compression for Long-Context LLM Inference
von: Shi, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Shi, Zhiyuan, et al.
Veröffentlicht: (2026)
SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention
von: Zhu, Qianchao, et al.
Veröffentlicht: (2024)
von: Zhu, Qianchao, et al.
Veröffentlicht: (2024)
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs
von: Yang, Dongquan, et al.
Veröffentlicht: (2025)
von: Yang, Dongquan, et al.
Veröffentlicht: (2025)
DepthKV: Layer-Dependent KV Cache Pruning for Long-Context LLM Inference
von: Dehghanighobadi, Zahra, et al.
Veröffentlicht: (2026)
von: Dehghanighobadi, Zahra, et al.
Veröffentlicht: (2026)
Reframing Tax Law Entailment as Analogical Reasoning
von: Zou, Xinrui, et al.
Veröffentlicht: (2024)
von: Zou, Xinrui, et al.
Veröffentlicht: (2024)
Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirements
von: Zhang, Jingyu, et al.
Veröffentlicht: (2024)
von: Zhang, Jingyu, et al.
Veröffentlicht: (2024)
Lossless Token Sequence Compression via Meta-Tokens
von: Harvill, John, et al.
Veröffentlicht: (2025)
von: Harvill, John, et al.
Veröffentlicht: (2025)
SurfaceLogicKV: Surface and Logic Attention Behaviors are All You Need for Robust KV Cache Compression
von: Li, Mengjie, et al.
Veröffentlicht: (2025)
von: Li, Mengjie, et al.
Veröffentlicht: (2025)
MatryoshkaKV: Adaptive KV Compression via Trainable Orthogonal Projection
von: Lin, Bokai, et al.
Veröffentlicht: (2024)
von: Lin, Bokai, et al.
Veröffentlicht: (2024)
KV Cache Steering for Controlling Frozen LLMs
von: Belitsky, Max, et al.
Veröffentlicht: (2025)
von: Belitsky, Max, et al.
Veröffentlicht: (2025)
Always Tell Me The Odds: Fine-grained Conditional Probability Estimation
von: Wang, Liaoyaqi, et al.
Veröffentlicht: (2025)
von: Wang, Liaoyaqi, et al.
Veröffentlicht: (2025)
CommVQ: Commutative Vector Quantization for KV Cache Compression
von: Li, Junyan, et al.
Veröffentlicht: (2025)
von: Li, Junyan, et al.
Veröffentlicht: (2025)
Leveraging KV Similarity for Online Structured Pruning in LLMs
von: Lee, Jungmin, et al.
Veröffentlicht: (2025)
von: Lee, Jungmin, et al.
Veröffentlicht: (2025)
LoRC: Low-Rank Compression for LLMs KV Cache with a Progressive Compression Strategy
von: Zhang, Rongzhi, et al.
Veröffentlicht: (2024)
von: Zhang, Rongzhi, et al.
Veröffentlicht: (2024)
KVReviver: Reversible KV Cache Compression with Sketch-Based Token Reconstruction
von: Yuan, Aomufei, et al.
Veröffentlicht: (2025)
von: Yuan, Aomufei, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Compactor: Calibrated Query-Agnostic KV Cache Compression with Approximate Leverage Scores
von: Chari, Vivek, et al.
Veröffentlicht: (2025) -
Lossless KV Cache Compression to 2%
von: Yang, Zhen, et al.
Veröffentlicht: (2024) -
GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
von: Kang, Hao, et al.
Veröffentlicht: (2024) -
KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
von: Tian, Yuxuan, et al.
Veröffentlicht: (2025) -
Dodo: Dynamic Contextual Compression for Decoder-only LMs
von: Qin, Guanghui, et al.
Veröffentlicht: (2023)