GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling
Fuente:
arXiv
Saved in:
| Main Authors: | Dadgarnia, Alireza, Tabesh, Soroush, Nikdan, Mahdi, Helcig, Michael, Kurtic, Eldar, Kleinegger, Maximilian, Alistarh, Dan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Statistically-Lossless Quantization of Large Language Models
by: Helcig, Michael, et al.
Published: (2026)
by: Helcig, Michael, et al.
Published: (2026)
RoSA: Accurate Parameter-Efficient Fine-Tuning via Robust Adaptation
by: Nikdan, Mahdi, et al.
Published: (2024)
by: Nikdan, Mahdi, et al.
Published: (2024)
HALO: Hadamard-Assisted Lower-Precision Optimization for LLMs
by: Ashkboos, Saleh, et al.
Published: (2025)
by: Ashkboos, Saleh, et al.
Published: (2025)
MatGPTQ: Accurate and Efficient Post-Training Matryoshka Quantization
by: Kleinegger, Maximilian, et al.
Published: (2026)
by: Kleinegger, Maximilian, et al.
Published: (2026)
QuEST: Stable Training of LLMs with 1-Bit Weights and Activations
by: Panferov, Andrei, et al.
Published: (2025)
by: Panferov, Andrei, et al.
Published: (2025)
ECO: Quantized Training without Full-Precision Master Weights
by: Nikdan, Mahdi, et al.
Published: (2026)
by: Nikdan, Mahdi, et al.
Published: (2026)
EvoPress: Accurate Dynamic Model Compression via Evolutionary Search
by: Sieberling, Oliver, et al.
Published: (2024)
by: Sieberling, Oliver, et al.
Published: (2024)
CAGE: Curvature-Aware Gradient Estimation For Accurate Quantization-Aware Training
by: Tabesh, Soroush, et al.
Published: (2025)
by: Tabesh, Soroush, et al.
Published: (2025)
Mathador-LM: A Dynamic Benchmark for Mathematical Reasoning on Large Language Models
by: Kurtic, Eldar, et al.
Published: (2024)
by: Kurtic, Eldar, et al.
Published: (2024)
Model Compression with Exact Budget Constraints via Riemannian Manifolds
by: Helcig, Michael, et al.
Published: (2026)
by: Helcig, Michael, et al.
Published: (2026)
Quartet II: Accurate LLM Pre-Training in NVFP4 by Improved Unbiased Gradient Estimation
by: Panferov, Andrei, et al.
Published: (2026)
by: Panferov, Andrei, et al.
Published: (2026)
Error Feedback Can Accurately Compress Preconditioners
by: Modoranu, Ionut-Vlad, et al.
Published: (2023)
by: Modoranu, Ionut-Vlad, et al.
Published: (2023)
"Give Me BF16 or Give Me Death"? Accuracy-Performance Trade-Offs in LLM Quantization
by: Kurtic, Eldar, et al.
Published: (2024)
by: Kurtic, Eldar, et al.
Published: (2024)
Panza: Design and Analysis of a Fully-Local Personalized Text Writing Assistant
by: Nicolicioiu, Armand, et al.
Published: (2024)
by: Nicolicioiu, Armand, et al.
Published: (2024)
Quartet: Native FP4 Training Can Be Optimal for Large Language Models
by: Castro, Roberto L., et al.
Published: (2025)
by: Castro, Roberto L., et al.
Published: (2025)
MicroAdam: Accurate Adaptive Optimization with Low Space Overhead and Provable Convergence
by: Modoranu, Ionut-Vlad, et al.
Published: (2024)
by: Modoranu, Ionut-Vlad, et al.
Published: (2024)
Efficient Data Selection at Scale via Influence Distillation
by: Nikdan, Mahdi, et al.
Published: (2025)
by: Nikdan, Mahdi, et al.
Published: (2025)
DarwinLM: Evolutionary Structured Pruning of Large Language Models
by: Tang, Shengkun, et al.
Published: (2025)
by: Tang, Shengkun, et al.
Published: (2025)
Gumbel-MPNN: Graph Rewiring with Gumbel-Softmax
by: Hoffmann, Marcel, et al.
Published: (2025)
by: Hoffmann, Marcel, et al.
Published: (2025)
Bridging the Gap Between Promise and Performance for Microscaling FP4 Quantization
by: Egiazarian, Vage, et al.
Published: (2025)
by: Egiazarian, Vage, et al.
Published: (2025)
GSQ-Tuning: Group-Shared Exponents Integer in Fully Quantized Training for LLMs On-Device Fine-tuning
by: Zhou, Sifan, et al.
Published: (2025)
by: Zhou, Sifan, et al.
Published: (2025)
Universal Adversarial Suffixes Using Calibrated Gumbel-Softmax Relaxation
by: Soor, Sampriti, et al.
Published: (2025)
by: Soor, Sampriti, et al.
Published: (2025)
LLMQ: Efficient Lower-Precision Pretraining for Consumer GPUs
by: Schultheis, Erik, et al.
Published: (2025)
by: Schultheis, Erik, et al.
Published: (2025)
Accurate Compression of Text-to-Image Diffusion Models via Vector Quantization
by: Egiazarian, Vage, et al.
Published: (2024)
by: Egiazarian, Vage, et al.
Published: (2024)
MatryoshkaLoRA: Learning Accurate Hierarchical Low-Rank Representations for LLM Fine-Tuning
by: Modoranu, Ionut-Vlad, et al.
Published: (2026)
by: Modoranu, Ionut-Vlad, et al.
Published: (2026)
LRW-Persian: Lip-reading in the Wild Dataset for Persian Language
by: Taghizadeh, Zahra, et al.
Published: (2025)
by: Taghizadeh, Zahra, et al.
Published: (2025)
Apertus LLM Family Expansion via Distillation and Quantization
by: Panferov, Andrei, et al.
Published: (2026)
by: Panferov, Andrei, et al.
Published: (2026)
Enabling High-Sparsity Foundational Llama Models with Efficient Pretraining and Deployment
by: Agarwalla, Abhinav, et al.
Published: (2024)
by: Agarwalla, Abhinav, et al.
Published: (2024)
Behemoth: Benchmarking Unlearning in LLMs Using Fully Synthetic Data
by: Iofinova, Eugenia, et al.
Published: (2026)
by: Iofinova, Eugenia, et al.
Published: (2026)
Federated Variational Preference Alignment with Gumbel-Softmax Prior for Personalized User Preferences
by: Koo, Jabin, et al.
Published: (2026)
by: Koo, Jabin, et al.
Published: (2026)
Joint Beamforming and Integer User Association using a GNN with Gumbel-Softmax Reparameterizations
by: Lyu, Qing, et al.
Published: (2025)
by: Lyu, Qing, et al.
Published: (2025)
Gumbel-Softmax Flow Matching with Straight-Through Guidance for Controllable Biological Sequence Generation
by: Tang, Sophia, et al.
Published: (2025)
by: Tang, Sophia, et al.
Published: (2025)
Can LLMs Separate Instructions From Data? And What Do We Even Mean By That?
by: Zverev, Egor, et al.
Published: (2024)
by: Zverev, Egor, et al.
Published: (2024)
Optimal sensor placement for the reconstruction of ocean states using differentiable Gumbel-Softmax sampling operator
by: Chapron, Oscar, et al.
Published: (2026)
by: Chapron, Oscar, et al.
Published: (2026)
Conditional Gumbel-Softmax for constrained feature selection with application to node selection in wireless sensor networks
by: Strypsteen, Thomas, et al.
Published: (2024)
by: Strypsteen, Thomas, et al.
Published: (2024)
SimA: Simple Softmax-free Attention for Vision Transformers
by: Koohpayegani, Soroush Abbasi, et al.
Published: (2022)
by: Koohpayegani, Soroush Abbasi, et al.
Published: (2022)
Speculative Decoding Speed-of-Light: Optimal Lower Bounds via Branching Random Walks
by: Pankratov, Sergey, et al.
Published: (2025)
by: Pankratov, Sergey, et al.
Published: (2025)
Extreme Compression of Large Language Models via Additive Quantization
by: Egiazarian, Vage, et al.
Published: (2024)
by: Egiazarian, Vage, et al.
Published: (2024)
Optimal design of frame structures with mixed categorical and continuous design variables using the Gumbel-Softmax method
by: Ebrahimi, Mehran, et al.
Published: (2024)
by: Ebrahimi, Mehran, et al.
Published: (2024)
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem
by: Malinovskii, Vladimir, et al.
Published: (2024)
by: Malinovskii, Vladimir, et al.
Published: (2024)
Similar Items
-
Statistically-Lossless Quantization of Large Language Models
by: Helcig, Michael, et al.
Published: (2026) -
RoSA: Accurate Parameter-Efficient Fine-Tuning via Robust Adaptation
by: Nikdan, Mahdi, et al.
Published: (2024) -
HALO: Hadamard-Assisted Lower-Precision Optimization for LLMs
by: Ashkboos, Saleh, et al.
Published: (2025) -
MatGPTQ: Accurate and Efficient Post-Training Matryoshka Quantization
by: Kleinegger, Maximilian, et al.
Published: (2026) -
QuEST: Stable Training of LLMs with 1-Bit Weights and Activations
by: Panferov, Andrei, et al.
Published: (2025)