Saved in:
| Main Authors: | Shutova, Alina, Malinovskii, Vladimir, Egiazarian, Vage, Kuznedelev, Denis, Mazur, Denis, Surkov, Nikita, Ermakov, Ivan, Alistarh, Dan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2501.19392 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Extreme Compression of Large Language Models via Additive Quantization
by: Egiazarian, Vage, et al.
Published: (2024)
by: Egiazarian, Vage, et al.
Published: (2024)
Hogwild! Inference: Parallel LLM Generation via Concurrent Attention
by: Rodionov, Gleb, et al.
Published: (2025)
by: Rodionov, Gleb, et al.
Published: (2025)
PV-Tuning: Beyond Straight-Through Estimation for Extreme LLM Compression
by: Malinovskii, Vladimir, et al.
Published: (2024)
by: Malinovskii, Vladimir, et al.
Published: (2024)
Accurate Compression of Text-to-Image Diffusion Models via Vector Quantization
by: Egiazarian, Vage, et al.
Published: (2024)
by: Egiazarian, Vage, et al.
Published: (2024)
WUSH: Near-Optimal Adaptive Transforms for LLM Quantization
by: Chen, Jiale, et al.
Published: (2025)
by: Chen, Jiale, et al.
Published: (2025)
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem
by: Malinovskii, Vladimir, et al.
Published: (2024)
by: Malinovskii, Vladimir, et al.
Published: (2024)
Grid Games: The Power of Multiple Grids for Quantizing Large Language Models
by: Egiazarian, Vage, et al.
Published: (2026)
by: Egiazarian, Vage, et al.
Published: (2026)
KV Cache Offloading for Context-Intensive Tasks
by: Bocharnikov, Andrey, et al.
Published: (2026)
by: Bocharnikov, Andrey, et al.
Published: (2026)
Bridging the Gap Between Promise and Performance for Microscaling FP4 Quantization
by: Egiazarian, Vage, et al.
Published: (2025)
by: Egiazarian, Vage, et al.
Published: (2025)
EvoPress: Accurate Dynamic Model Compression via Evolutionary Search
by: Sieberling, Oliver, et al.
Published: (2024)
by: Sieberling, Oliver, et al.
Published: (2024)
AutoJudge: Judge Decoding Without Manual Annotation
by: Garipov, Roman, et al.
Published: (2025)
by: Garipov, Roman, et al.
Published: (2025)
The Iterative Optimal Brain Surgeon: Faster Sparse Recovery by Leveraging Second-Order Information
by: Wu, Diyuan, et al.
Published: (2024)
by: Wu, Diyuan, et al.
Published: (2024)
Unified Scaling Laws for Compressed Representations
by: Panferov, Andrei, et al.
Published: (2025)
by: Panferov, Andrei, et al.
Published: (2025)
Asynchronous Reasoning: Training-Free Interactive Thinking LLMs
by: Yakushev, George, et al.
Published: (2025)
by: Yakushev, George, et al.
Published: (2025)
Statistically-Lossless Quantization of Large Language Models
by: Helcig, Michael, et al.
Published: (2026)
by: Helcig, Michael, et al.
Published: (2026)
"Give Me BF16 or Give Me Death"? Accuracy-Performance Trade-Offs in LLM Quantization
by: Kurtic, Eldar, et al.
Published: (2024)
by: Kurtic, Eldar, et al.
Published: (2024)
Scale-wise Distillation of Diffusion Models
by: Starodubcev, Nikita, et al.
Published: (2025)
by: Starodubcev, Nikita, et al.
Published: (2025)
SKVQ: Sliding-window Key and Value Cache Quantization for Large Language Models
by: Duanmu, Haojie, et al.
Published: (2024)
by: Duanmu, Haojie, et al.
Published: (2024)
WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More
by: Yue, Yuxuan, et al.
Published: (2024)
by: Yue, Yuxuan, et al.
Published: (2024)
Refined distributional limit theorems for compound sums
by: Malinovskii, Vsevolod K.
Published: (2024)
by: Malinovskii, Vsevolod K.
Published: (2024)
Does Diffusion Beat GAN in Image Super Resolution?
by: Kuznedelev, Denis, et al.
Published: (2024)
by: Kuznedelev, Denis, et al.
Published: (2024)
Characterizing Graph Datasets for Node Classification: Homophily-Heterophily Dichotomy and Beyond
by: Platonov, Oleg, et al.
Published: (2022)
by: Platonov, Oleg, et al.
Published: (2022)
MatGPTQ: Accurate and Efficient Post-Training Matryoshka Quantization
by: Kleinegger, Maximilian, et al.
Published: (2026)
by: Kleinegger, Maximilian, et al.
Published: (2026)
Evaluating Memory Structure in LLM Agents
by: Shutova, Alina, et al.
Published: (2026)
by: Shutova, Alina, et al.
Published: (2026)
Neural Optimal Transport with General Cost Functionals
by: Asadulaev, Arip, et al.
Published: (2022)
by: Asadulaev, Arip, et al.
Published: (2022)
Label Privacy in Split Learning for Large Models with Parameter-Efficient Training
by: Zmushko, Philip, et al.
Published: (2024)
by: Zmushko, Philip, et al.
Published: (2024)
Mathador-LM: A Dynamic Benchmark for Mathematical Reasoning on Large Language Models
by: Kurtic, Eldar, et al.
Published: (2024)
by: Kurtic, Eldar, et al.
Published: (2024)
Apertus LLM Family Expansion via Distillation and Quantization
by: Panferov, Andrei, et al.
Published: (2026)
by: Panferov, Andrei, et al.
Published: (2026)
ECO: Quantized Training without Full-Precision Master Weights
by: Nikdan, Mahdi, et al.
Published: (2026)
by: Nikdan, Mahdi, et al.
Published: (2026)
Switti: Designing Scale-Wise Transformers for Text-to-Image Synthesis
by: Voronov, Anton, et al.
Published: (2024)
by: Voronov, Anton, et al.
Published: (2024)
A critical look at the evaluation of GNNs under heterophily: Are we really making progress?
by: Platonov, Oleg, et al.
Published: (2023)
by: Platonov, Oleg, et al.
Published: (2023)
Self-Alignment: Improving Alignment of Cultural Values in LLMs via In-Context Learning
by: Choenni, Rochelle, et al.
Published: (2024)
by: Choenni, Rochelle, et al.
Published: (2024)
Rethinking Optimal Transport in Offline Reinforcement Learning
by: Asadulaev, Arip, et al.
Published: (2024)
by: Asadulaev, Arip, et al.
Published: (2024)
Mitigating the Impact of Outlier Channels for Language Model Quantization with Activation Regularization
by: Nrusimha, Aniruddha, et al.
Published: (2024)
by: Nrusimha, Aniruddha, et al.
Published: (2024)
Adaptive Cost Model for Query Optimization
by: Vasilenko, Nikita, et al.
Published: (2024)
by: Vasilenko, Nikita, et al.
Published: (2024)
Cache Me if You Can: Accelerating Diffusion Models through Block Caching
by: Wimbauer, Felix, et al.
Published: (2023)
by: Wimbauer, Felix, et al.
Published: (2023)
Rethinking Key-Value Cache Compression Techniques for Large Language Model Serving
by: Gao, Wei, et al.
Published: (2025)
by: Gao, Wei, et al.
Published: (2025)
Compression Scaling Laws:Unifying Sparsity and Quantization
by: Frantar, Elias, et al.
Published: (2025)
by: Frantar, Elias, et al.
Published: (2025)
Are LLMs classical or nonmonotonic reasoners? Lessons from generics
by: Leidinger, Alina, et al.
Published: (2024)
by: Leidinger, Alina, et al.
Published: (2024)
Beyond Outliers: A Study of Optimizers Under Quantization
by: Vlassis, Georgios, et al.
Published: (2025)
by: Vlassis, Georgios, et al.
Published: (2025)
Similar Items
-
Extreme Compression of Large Language Models via Additive Quantization
by: Egiazarian, Vage, et al.
Published: (2024) -
Hogwild! Inference: Parallel LLM Generation via Concurrent Attention
by: Rodionov, Gleb, et al.
Published: (2025) -
PV-Tuning: Beyond Straight-Through Estimation for Extreme LLM Compression
by: Malinovskii, Vladimir, et al.
Published: (2024) -
Accurate Compression of Text-to-Image Diffusion Models via Vector Quantization
by: Egiazarian, Vage, et al.
Published: (2024) -
WUSH: Near-Optimal Adaptive Transforms for LLM Quantization
by: Chen, Jiale, et al.
Published: (2025)