Polynomial Context-Truncation Sensitivity in Autoregressive Language Models: Sequential Wyner-Ziv Bounds for KV Cache Compression
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Kim, Munsik |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FibQuant: Universal Vector Quantization for Random-Access KV-Cache Compression
von: Lee, Namyoon, et al.
Veröffentlicht: (2026)
von: Lee, Namyoon, et al.
Veröffentlicht: (2026)
Bottlenecked Transformers: Periodic KV Cache Consolidation for Generalised Reasoning
von: Oomerjee, Adnan, et al.
Veröffentlicht: (2025)
von: Oomerjee, Adnan, et al.
Veröffentlicht: (2025)
Rethinking KV Cache Eviction via a Unified Information-Theoretic Objective
von: Yang, Jiaming, et al.
Veröffentlicht: (2026)
von: Yang, Jiaming, et al.
Veröffentlicht: (2026)
SQuat: Subspace-orthogonal KV Cache Quantization
von: Wang, Hao, et al.
Veröffentlicht: (2025)
von: Wang, Hao, et al.
Veröffentlicht: (2025)
One-Shot Wyner-Ziv Compression of a Uniform Source
von: Ülger, Oğuzhan Kubilay, et al.
Veröffentlicht: (2024)
von: Ülger, Oğuzhan Kubilay, et al.
Veröffentlicht: (2024)
Sequential KV Cache Compression via Probabilistic Language Tries: Beyond the Per-Vector Shannon Limit
von: Magarshak, Gregory
Veröffentlicht: (2026)
von: Magarshak, Gregory
Veröffentlicht: (2026)
Language Modeling Is Compression
von: Delétang, Grégoire, et al.
Veröffentlicht: (2023)
von: Delétang, Grégoire, et al.
Veröffentlicht: (2023)
The Pitfalls of KV Cache Compression
von: Chen, Alex, et al.
Veröffentlicht: (2025)
von: Chen, Alex, et al.
Veröffentlicht: (2025)
Beyond RAG: Task-Aware KV Cache Compression for Comprehensive Knowledge Reasoning
von: Corallo, Giulio, et al.
Veröffentlicht: (2025)
von: Corallo, Giulio, et al.
Veröffentlicht: (2025)
Informationally Compressive Anonymization: Non-Degrading Sensitive Input Protection for Privacy-Preserving Supervised Machine Learning
von: Samuelson, Jeremy J
Veröffentlicht: (2026)
von: Samuelson, Jeremy J
Veröffentlicht: (2026)
OjaKV: Context-Aware Online Low-Rank KV Cache Compression
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
Contextual Control without Memory Growth in a Context-Switching Task
von: Kim, Song-Ju
Veröffentlicht: (2026)
von: Kim, Song-Ju
Veröffentlicht: (2026)
MiniCache: KV Cache Compression in Depth Dimension for Large Language Models
von: Liu, Akide, et al.
Veröffentlicht: (2024)
von: Liu, Akide, et al.
Veröffentlicht: (2024)
CommonKV: Compressing KV Cache with Cross-layer Parameter Sharing
von: Wang, Yixuan, et al.
Veröffentlicht: (2025)
von: Wang, Yixuan, et al.
Veröffentlicht: (2025)
Scalar Lattices and Probabilistic Shaping for Dithered Wyner-Ziv Quantization
von: Sener, Muhammed Yusuf, et al.
Veröffentlicht: (2025)
von: Sener, Muhammed Yusuf, et al.
Veröffentlicht: (2025)
SmallKV: Small Model Assisted Compensation of KV Cache Compression for Efficient LLM Inference
von: Zhao, Yi, et al.
Veröffentlicht: (2025)
von: Zhao, Yi, et al.
Veröffentlicht: (2025)
KV Packet: Recomputation-Free Context-Independent KV Caching for LLMs
von: Chen, Chuangtao, et al.
Veröffentlicht: (2026)
von: Chen, Chuangtao, et al.
Veröffentlicht: (2026)
In-Context Learning for MIMO Equalization Using Transformer-Based Sequence Models
von: Zecchin, Matteo, et al.
Veröffentlicht: (2023)
von: Zecchin, Matteo, et al.
Veröffentlicht: (2023)
Compressing Chemistry Reveals Functional Groups
von: Sharma, Ruben, et al.
Veröffentlicht: (2025)
von: Sharma, Ruben, et al.
Veröffentlicht: (2025)
Distributed and Rate-Adaptive Feature Compression
von: Deshmukh, Aditya, et al.
Veröffentlicht: (2024)
von: Deshmukh, Aditya, et al.
Veröffentlicht: (2024)
Hurwitz Quaternion Multiplicative Quantization for KV Cache Compression
von: Swain, Kabir, et al.
Veröffentlicht: (2026)
von: Swain, Kabir, et al.
Veröffentlicht: (2026)
Palu: Compressing KV-Cache with Low-Rank Projection
von: Chang, Chi-Chih, et al.
Veröffentlicht: (2024)
von: Chang, Chi-Chih, et al.
Veröffentlicht: (2024)
KVSculpt: KV Cache Compression as Distillation
von: Jiang, Bo, et al.
Veröffentlicht: (2026)
von: Jiang, Bo, et al.
Veröffentlicht: (2026)
Enhancing Large Multimodal Models with Adaptive Sparsity and KV Cache Compression
von: Zhang, Te, et al.
Veröffentlicht: (2025)
von: Zhang, Te, et al.
Veröffentlicht: (2025)
Cache What Lasts: Token Retention for Memory-Bounded KV Cache in LLMs
von: Bui, Ngoc, et al.
Veröffentlicht: (2025)
von: Bui, Ngoc, et al.
Veröffentlicht: (2025)
Training-free Adjustable Polynomial Graph Filtering for Ultra-fast Multimodal Recommendation
von: Roh, Yu-Seung, et al.
Veröffentlicht: (2025)
von: Roh, Yu-Seung, et al.
Veröffentlicht: (2025)
Optimizing Learned Image Compression on Scalar and Entropy-Constraint Quantization
von: Borzechowski, Florian, et al.
Veröffentlicht: (2025)
von: Borzechowski, Florian, et al.
Veröffentlicht: (2025)
AlphaZip: Neural Network-Enhanced Lossless Text Compression
von: Narashiman, Swathi Shree, et al.
Veröffentlicht: (2024)
von: Narashiman, Swathi Shree, et al.
Veröffentlicht: (2024)
LaCache: Ladder-Shaped KV Caching for Efficient Long-Context Modeling of Large Language Models
von: Shi, Dachuan, et al.
Veröffentlicht: (2025)
von: Shi, Dachuan, et al.
Veröffentlicht: (2025)
The Illusion of Equivalence: Systematic FP16 Divergence in KV-Cached Autoregressive Inference
von: Chodavarapu, Ranjith, et al.
Veröffentlicht: (2026)
von: Chodavarapu, Ranjith, et al.
Veröffentlicht: (2026)
A General Error-Theoretical Analysis Framework for Constructing Compression Strategies
von: Zhang, Boyang, et al.
Veröffentlicht: (2025)
von: Zhang, Boyang, et al.
Veröffentlicht: (2025)
Interpretability as Compression: Reconsidering SAE Explanations of Neural Activations with MDL-SAEs
von: Ayonrinde, Kola, et al.
Veröffentlicht: (2024)
von: Ayonrinde, Kola, et al.
Veröffentlicht: (2024)
ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression
von: Liu, Guangda, et al.
Veröffentlicht: (2024)
von: Liu, Guangda, et al.
Veröffentlicht: (2024)
Near-Oracle KV Selection via Pre-hoc Sparsity for Long-Context Inference
von: Gao, Yifei, et al.
Veröffentlicht: (2026)
von: Gao, Yifei, et al.
Veröffentlicht: (2026)
GeoIB: Geometry-Aware Information Bottleneck via Statistical-Manifold Compression
von: Wang, Weiqi, et al.
Veröffentlicht: (2026)
von: Wang, Weiqi, et al.
Veröffentlicht: (2026)
Haiku to Opus in Just 10 bits: LLMs Unlock Massive Compression Gains
von: Rinberg, Roy, et al.
Veröffentlicht: (2026)
von: Rinberg, Roy, et al.
Veröffentlicht: (2026)
Flexible Variational Information Bottleneck: Achieving Diverse Compression with a Single Training
von: Kudo, Sota, et al.
Veröffentlicht: (2024)
von: Kudo, Sota, et al.
Veröffentlicht: (2024)
When Wyner and Ziv Met Bayes in Quantum-Classical Realm
von: Sohail, Mohammad Aamir, et al.
Veröffentlicht: (2025)
von: Sohail, Mohammad Aamir, et al.
Veröffentlicht: (2025)
Understanding LLM Behaviors via Compression: Data Generation, Knowledge Acquisition and Scaling Laws
von: Pan, Zhixuan, et al.
Veröffentlicht: (2025)
von: Pan, Zhixuan, et al.
Veröffentlicht: (2025)
Compression via Pre-trained Transformers: A Study on Byte-Level Multimodal Data
von: Heurtel-Depeiges, David, et al.
Veröffentlicht: (2024)
von: Heurtel-Depeiges, David, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
FibQuant: Universal Vector Quantization for Random-Access KV-Cache Compression
von: Lee, Namyoon, et al.
Veröffentlicht: (2026) -
Bottlenecked Transformers: Periodic KV Cache Consolidation for Generalised Reasoning
von: Oomerjee, Adnan, et al.
Veröffentlicht: (2025) -
Rethinking KV Cache Eviction via a Unified Information-Theoretic Objective
von: Yang, Jiaming, et al.
Veröffentlicht: (2026) -
SQuat: Subspace-orthogonal KV Cache Quantization
von: Wang, Hao, et al.
Veröffentlicht: (2025) -
One-Shot Wyner-Ziv Compression of a Uniform Source
von: Ülger, Oğuzhan Kubilay, et al.
Veröffentlicht: (2024)