MatGPTQ: Accurate and Efficient Post-Training Matryoshka Quantization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kleinegger, Maximilian, Crnčević, Elvir, Alistarh, Dan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Geometry of LLM Quantization: GPTQ as Babai's Nearest Plane Algorithm
von: Chen, Jiale, et al.
Veröffentlicht: (2025)
von: Chen, Jiale, et al.
Veröffentlicht: (2025)
RoSA: Accurate Parameter-Efficient Fine-Tuning via Robust Adaptation
von: Nikdan, Mahdi, et al.
Veröffentlicht: (2024)
von: Nikdan, Mahdi, et al.
Veröffentlicht: (2024)
GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling
von: Dadgarnia, Alireza, et al.
Veröffentlicht: (2026)
von: Dadgarnia, Alireza, et al.
Veröffentlicht: (2026)
MatryoshkaLoRA: Learning Accurate Hierarchical Low-Rank Representations for LLM Fine-Tuning
von: Modoranu, Ionut-Vlad, et al.
Veröffentlicht: (2026)
von: Modoranu, Ionut-Vlad, et al.
Veröffentlicht: (2026)
CAGE: Curvature-Aware Gradient Estimation For Accurate Quantization-Aware Training
von: Tabesh, Soroush, et al.
Veröffentlicht: (2025)
von: Tabesh, Soroush, et al.
Veröffentlicht: (2025)
Matryoshka Quantization
von: Nair, Pranav, et al.
Veröffentlicht: (2025)
von: Nair, Pranav, et al.
Veröffentlicht: (2025)
Quartet II: Accurate LLM Pre-Training in NVFP4 by Improved Unbiased Gradient Estimation
von: Panferov, Andrei, et al.
Veröffentlicht: (2026)
von: Panferov, Andrei, et al.
Veröffentlicht: (2026)
Statistically-Lossless Quantization of Large Language Models
von: Helcig, Michael, et al.
Veröffentlicht: (2026)
von: Helcig, Michael, et al.
Veröffentlicht: (2026)
ECO: Quantized Training without Full-Precision Master Weights
von: Nikdan, Mahdi, et al.
Veröffentlicht: (2026)
von: Nikdan, Mahdi, et al.
Veröffentlicht: (2026)
MatMamba: A Matryoshka State Space Model
von: Shukla, Abhinav, et al.
Veröffentlicht: (2024)
von: Shukla, Abhinav, et al.
Veröffentlicht: (2024)
Apertus LLM Family Expansion via Distillation and Quantization
von: Panferov, Andrei, et al.
Veröffentlicht: (2026)
von: Panferov, Andrei, et al.
Veröffentlicht: (2026)
EvoPress: Accurate Dynamic Model Compression via Evolutionary Search
von: Sieberling, Oliver, et al.
Veröffentlicht: (2024)
von: Sieberling, Oliver, et al.
Veröffentlicht: (2024)
SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models
von: Xiao, Guangxuan, et al.
Veröffentlicht: (2022)
von: Xiao, Guangxuan, et al.
Veröffentlicht: (2022)
LLMQ: Efficient Lower-Precision Pretraining for Consumer GPUs
von: Schultheis, Erik, et al.
Veröffentlicht: (2025)
von: Schultheis, Erik, et al.
Veröffentlicht: (2025)
Beyond Outliers: A Study of Optimizers Under Quantization
von: Vlassis, Georgios, et al.
Veröffentlicht: (2025)
von: Vlassis, Georgios, et al.
Veröffentlicht: (2025)
WUSH: Near-Optimal Adaptive Transforms for LLM Quantization
von: Chen, Jiale, et al.
Veröffentlicht: (2025)
von: Chen, Jiale, et al.
Veröffentlicht: (2025)
Behemoth: Benchmarking Unlearning in LLMs Using Fully Synthetic Data
von: Iofinova, Eugenia, et al.
Veröffentlicht: (2026)
von: Iofinova, Eugenia, et al.
Veröffentlicht: (2026)
Model Compression with Exact Budget Constraints via Riemannian Manifolds
von: Helcig, Michael, et al.
Veröffentlicht: (2026)
von: Helcig, Michael, et al.
Veröffentlicht: (2026)
Compression Scaling Laws:Unifying Sparsity and Quantization
von: Frantar, Elias, et al.
Veröffentlicht: (2025)
von: Frantar, Elias, et al.
Veröffentlicht: (2025)
GPTQ-intrinsic LoRA: A Near-optimal Algorithm for Low-precision Quantization with Low-rank Adaptation
von: Zhang, Shihao, et al.
Veröffentlicht: (2026)
von: Zhang, Shihao, et al.
Veröffentlicht: (2026)
The Lattice Geometry of Neural Network Quantization -- A Short Equivalence Proof of GPTQ and Babai's Algorithm
von: Birnick, Johann
Veröffentlicht: (2025)
von: Birnick, Johann
Veröffentlicht: (2025)
Communication-Efficient Federated Learning With Data and Client Heterogeneity
von: Zakerinia, Hossein, et al.
Veröffentlicht: (2022)
von: Zakerinia, Hossein, et al.
Veröffentlicht: (2022)
Grid Games: The Power of Multiple Grids for Quantizing Large Language Models
von: Egiazarian, Vage, et al.
Veröffentlicht: (2026)
von: Egiazarian, Vage, et al.
Veröffentlicht: (2026)
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem
von: Malinovskii, Vladimir, et al.
Veröffentlicht: (2024)
von: Malinovskii, Vladimir, et al.
Veröffentlicht: (2024)
D$^2$Quant: Accurate Low-bit Post-Training Weight Quantization for LLMs
von: Yan, Xianglong, et al.
Veröffentlicht: (2026)
von: Yan, Xianglong, et al.
Veröffentlicht: (2026)
RepQuant: Towards Accurate Post-Training Quantization of Large Transformer Models via Scale Reparameterization
von: Li, Zhikai, et al.
Veröffentlicht: (2024)
von: Li, Zhikai, et al.
Veröffentlicht: (2024)
Error Feedback Can Accurately Compress Preconditioners
von: Modoranu, Ionut-Vlad, et al.
Veröffentlicht: (2023)
von: Modoranu, Ionut-Vlad, et al.
Veröffentlicht: (2023)
"Give Me BF16 or Give Me Death"? Accuracy-Performance Trade-Offs in LLM Quantization
von: Kurtic, Eldar, et al.
Veröffentlicht: (2024)
von: Kurtic, Eldar, et al.
Veröffentlicht: (2024)
Extreme Compression of Large Language Models via Additive Quantization
von: Egiazarian, Vage, et al.
Veröffentlicht: (2024)
von: Egiazarian, Vage, et al.
Veröffentlicht: (2024)
Layer-wise Quantization for Quantized Optimistic Dual Averaging
von: Nguyen, Anh Duc, et al.
Veröffentlicht: (2025)
von: Nguyen, Anh Duc, et al.
Veröffentlicht: (2025)
Mitigating the Impact of Outlier Channels for Language Model Quantization with Activation Regularization
von: Nrusimha, Aniruddha, et al.
Veröffentlicht: (2024)
von: Nrusimha, Aniruddha, et al.
Veröffentlicht: (2024)
Training Dynamics Impact Post-Training Quantization Robustness
von: Catalan-Tatjer, Albert, et al.
Veröffentlicht: (2025)
von: Catalan-Tatjer, Albert, et al.
Veröffentlicht: (2025)
Powerset Convolutional Neural Networks
von: Wendler, Chris, et al.
Veröffentlicht: (2019)
von: Wendler, Chris, et al.
Veröffentlicht: (2019)
SKIM: Any-bit Quantization Pushing The Limits of Post-Training Quantization
von: Bai, Runsheng, et al.
Veröffentlicht: (2024)
von: Bai, Runsheng, et al.
Veröffentlicht: (2024)
Efficient Data Selection at Scale via Influence Distillation
von: Nikdan, Mahdi, et al.
Veröffentlicht: (2025)
von: Nikdan, Mahdi, et al.
Veröffentlicht: (2025)
Matryoshka Concept Bottleneck Models
von: Chen, Ziye, et al.
Veröffentlicht: (2026)
von: Chen, Ziye, et al.
Veröffentlicht: (2026)
OAC: Output-adaptive Calibration for Accurate Post-training Quantization
von: Edalati, Ali, et al.
Veröffentlicht: (2024)
von: Edalati, Ali, et al.
Veröffentlicht: (2024)
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models
von: Shutova, Alina, et al.
Veröffentlicht: (2025)
von: Shutova, Alina, et al.
Veröffentlicht: (2025)
Data Generation for Hardware-Friendly Post-Training Quantization
von: Dikstein, Lior, et al.
Veröffentlicht: (2024)
von: Dikstein, Lior, et al.
Veröffentlicht: (2024)
DASH: Faster Shampoo via Batched Block Preconditioning and Efficient Inverse-Root Solvers
von: Modoranu, Ionut-Vlad, et al.
Veröffentlicht: (2026)
von: Modoranu, Ionut-Vlad, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
The Geometry of LLM Quantization: GPTQ as Babai's Nearest Plane Algorithm
von: Chen, Jiale, et al.
Veröffentlicht: (2025) -
RoSA: Accurate Parameter-Efficient Fine-Tuning via Robust Adaptation
von: Nikdan, Mahdi, et al.
Veröffentlicht: (2024) -
GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling
von: Dadgarnia, Alireza, et al.
Veröffentlicht: (2026) -
MatryoshkaLoRA: Learning Accurate Hierarchical Low-Rank Representations for LLM Fine-Tuning
von: Modoranu, Ionut-Vlad, et al.
Veröffentlicht: (2026) -
CAGE: Curvature-Aware Gradient Estimation For Accurate Quantization-Aware Training
von: Tabesh, Soroush, et al.
Veröffentlicht: (2025)