Training Transformers for KV Cache Compressibility
Fuente:
arXiv
Saved in:
| Main Authors: | Gelberg, Yoav, Eitan, Yam, Bronstein, Michael, Gal, Yarin, Maron, Haggai |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Topological Blindspots: Understanding and Extending Topological Deep Learning Through the Lens of Expressivity
by: Eitan, Yam, et al.
Published: (2024)
by: Eitan, Yam, et al.
Published: (2024)
GradMetaNet: An Equivariant Architecture for Learning on Gradients
by: Gelberg, Yoav, et al.
Published: (2025)
by: Gelberg, Yoav, et al.
Published: (2025)
On The Expressive Power of GNN Derivatives
by: Eitan, Yam, et al.
Published: (2025)
by: Eitan, Yam, et al.
Published: (2025)
On the Expressive Power of Permutation-Equivariant Weight-Space Networks
by: Dayan, Adir, et al.
Published: (2026)
by: Dayan, Adir, et al.
Published: (2026)
Learning on LoRAs: GL-Equivariant Processing of Low-Rank Weight Spaces for Large Finetuned Models
by: Putterman, Theo, et al.
Published: (2024)
by: Putterman, Theo, et al.
Published: (2024)
Balancing Efficiency and Expressiveness: Subgraph GNNs with Walk-Based Centrality
by: Southern, Joshua, et al.
Published: (2025)
by: Southern, Joshua, et al.
Published: (2025)
A Flexible, Equivariant Framework for Subgraph GNNs via Graph Products and Graph Coarsening
by: Bar-Shalom, Guy, et al.
Published: (2024)
by: Bar-Shalom, Guy, et al.
Published: (2024)
FS-KAN: Permutation Equivariant Kolmogorov-Arnold Networks via Function Sharing
by: Elbaz, Ran, et al.
Published: (2025)
by: Elbaz, Ran, et al.
Published: (2025)
Beyond Next Token Probabilities: Learnable, Fast Detection of Hallucinations and Data Contamination on LLM Output Distributions
by: Bar-Shalom, Guy, et al.
Published: (2025)
by: Bar-Shalom, Guy, et al.
Published: (2025)
Efficient GNN Training Through Structure-Aware Randomized Mini-Batching
by: Balaji, Vignesh, et al.
Published: (2025)
by: Balaji, Vignesh, et al.
Published: (2025)
Variational Inference Failures Under Model Symmetries: Permutation Invariant Posteriors for Bayesian Neural Networks
by: Gelberg, Yoav, et al.
Published: (2024)
by: Gelberg, Yoav, et al.
Published: (2024)
Iterative Deployment Improves Planning Skills in LLMs
by: Corrêa, Augusto B., et al.
Published: (2025)
by: Corrêa, Augusto B., et al.
Published: (2025)
Subgraphormer: Unifying Subgraph GNNs and Graph Transformers via Graph Products
by: Bar-Shalom, Guy, et al.
Published: (2024)
by: Bar-Shalom, Guy, et al.
Published: (2024)
On the Reconstruction of Training Data from Group Invariant Networks
by: Elbaz, Ran, et al.
Published: (2024)
by: Elbaz, Ran, et al.
Published: (2024)
Foldable SuperNets: Scalable Merging of Transformers with Different Initializations and Tasks
by: Kinderman, Edan, et al.
Published: (2024)
by: Kinderman, Edan, et al.
Published: (2024)
The Pitfalls of KV Cache Compression
by: Chen, Alex, et al.
Published: (2025)
by: Chen, Alex, et al.
Published: (2025)
It Takes a Graph to Know a Graph: Rewiring for Homophily with a Reference Graph
by: Mendelman, Harel, et al.
Published: (2025)
by: Mendelman, Harel, et al.
Published: (2025)
Spanning the Visual Analogy Space with a Weight Basis of LoRAs
by: Manor, Hila, et al.
Published: (2026)
by: Manor, Hila, et al.
Published: (2026)
Equivariant Deep Weight Space Alignment
by: Navon, Aviv, et al.
Published: (2023)
by: Navon, Aviv, et al.
Published: (2023)
ManifoldKV: Training-Free KV Cache Compression via Euclidean Outlier Detection
by: Datta, Debajyoti, et al.
Published: (2026)
by: Datta, Debajyoti, et al.
Published: (2026)
KV-CAR: KV Cache Compression using Autoencoders and KV Reuse in Large Language Models
by: Roy, Sourjya, et al.
Published: (2025)
by: Roy, Sourjya, et al.
Published: (2025)
ThinKV: Thought-Adaptive KV Cache Compression for Efficient Reasoning Models
by: Ramachandran, Akshat, et al.
Published: (2025)
by: Ramachandran, Akshat, et al.
Published: (2025)
KVSculpt: KV Cache Compression as Distillation
by: Jiang, Bo, et al.
Published: (2026)
by: Jiang, Bo, et al.
Published: (2026)
CommonKV: Compressing KV Cache with Cross-layer Parameter Sharing
by: Wang, Yixuan, et al.
Published: (2025)
by: Wang, Yixuan, et al.
Published: (2025)
How Much Cache Does Reasoning Need? Depth-Cache Tradeoffs in KV-Compressed Transformers
by: Wang, Xiao
Published: (2026)
by: Wang, Xiao
Published: (2026)
EvolKV: Evolutionary KV Cache Compression for LLM Inference
by: Yu, Bohan, et al.
Published: (2025)
by: Yu, Bohan, et al.
Published: (2025)
KVCompose: Efficient Structured KV Cache Compression with Composite Tokens
by: Akulov, Dmitry, et al.
Published: (2025)
by: Akulov, Dmitry, et al.
Published: (2025)
Muon is Not That Special: Random or Inverted Spectra Work Just as Well
by: Shumaylov, Zakhar, et al.
Published: (2026)
by: Shumaylov, Zakhar, et al.
Published: (2026)
KV Pareto: Systems-Level Optimization of KV Cache and Model Compression for Long Context Inference
by: Gokhale, Sai, et al.
Published: (2025)
by: Gokhale, Sai, et al.
Published: (2025)
Hurwitz Quaternion Multiplicative Quantization for KV Cache Compression
by: Swain, Kabir, et al.
Published: (2026)
by: Swain, Kabir, et al.
Published: (2026)
Inference-Time Hyper-Scaling with KV Cache Compression
by: Łańcucki, Adrian, et al.
Published: (2025)
by: Łańcucki, Adrian, et al.
Published: (2025)
Palu: Compressing KV-Cache with Low-Rank Projection
by: Chang, Chi-Chih, et al.
Published: (2024)
by: Chang, Chi-Chih, et al.
Published: (2024)
ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression
by: Liu, Guangda, et al.
Published: (2024)
by: Liu, Guangda, et al.
Published: (2024)
KVzip: Query-Agnostic KV Cache Compression with Context Reconstruction
by: Kim, Jang-Hyun, et al.
Published: (2025)
by: Kim, Jang-Hyun, et al.
Published: (2025)
FastKV: Decoupling of Context Reduction and KV Cache Compression for Prefill-Decoding Acceleration
by: Jo, Dongwon, et al.
Published: (2025)
by: Jo, Dongwon, et al.
Published: (2025)
CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving
by: Liu, Yuhan, et al.
Published: (2023)
by: Liu, Yuhan, et al.
Published: (2023)
A Simple Plug-in for Improving Eviction-Based KV Cache Compression
by: Lin, Yuping, et al.
Published: (2026)
by: Lin, Yuping, et al.
Published: (2026)
KaVa: Latent Reasoning via Compressed KV-Cache Distillation
by: Kuzina, Anna, et al.
Published: (2025)
by: Kuzina, Anna, et al.
Published: (2025)
KQ-SVD: Compressing the KV Cache with Provable Guarantees on Attention Fidelity
by: Lesens, Damien, et al.
Published: (2025)
by: Lesens, Damien, et al.
Published: (2025)
Neural Message-Passing on Attention Graphs for Hallucination Detection
by: Frasca, Fabrizio, et al.
Published: (2025)
by: Frasca, Fabrizio, et al.
Published: (2025)
Similar Items
-
Topological Blindspots: Understanding and Extending Topological Deep Learning Through the Lens of Expressivity
by: Eitan, Yam, et al.
Published: (2024) -
GradMetaNet: An Equivariant Architecture for Learning on Gradients
by: Gelberg, Yoav, et al.
Published: (2025) -
On The Expressive Power of GNN Derivatives
by: Eitan, Yam, et al.
Published: (2025) -
On the Expressive Power of Permutation-Equivariant Weight-Space Networks
by: Dayan, Adir, et al.
Published: (2026) -
Learning on LoRAs: GL-Equivariant Processing of Low-Rank Weight Spaces for Large Finetuned Models
by: Putterman, Theo, et al.
Published: (2024)