Grid Games: The Power of Multiple Grids for Quantizing Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Egiazarian, Vage, Schultheis, Erik, Panferov, Andrei, Killian, Earl, Hoefler, Torsten, Alistarh, Dan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Extreme Compression of Large Language Models via Additive Quantization
von: Egiazarian, Vage, et al.
Veröffentlicht: (2024)
von: Egiazarian, Vage, et al.
Veröffentlicht: (2024)
WUSH: Near-Optimal Adaptive Transforms for LLM Quantization
von: Chen, Jiale, et al.
Veröffentlicht: (2025)
von: Chen, Jiale, et al.
Veröffentlicht: (2025)
Quartet II: Accurate LLM Pre-Training in NVFP4 by Improved Unbiased Gradient Estimation
von: Panferov, Andrei, et al.
Veröffentlicht: (2026)
von: Panferov, Andrei, et al.
Veröffentlicht: (2026)
Bridging the Gap Between Promise and Performance for Microscaling FP4 Quantization
von: Egiazarian, Vage, et al.
Veröffentlicht: (2025)
von: Egiazarian, Vage, et al.
Veröffentlicht: (2025)
Unified Scaling Laws for Compressed Representations
von: Panferov, Andrei, et al.
Veröffentlicht: (2025)
von: Panferov, Andrei, et al.
Veröffentlicht: (2025)
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem
von: Malinovskii, Vladimir, et al.
Veröffentlicht: (2024)
von: Malinovskii, Vladimir, et al.
Veröffentlicht: (2024)
Apertus LLM Family Expansion via Distillation and Quantization
von: Panferov, Andrei, et al.
Veröffentlicht: (2026)
von: Panferov, Andrei, et al.
Veröffentlicht: (2026)
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models
von: Shutova, Alina, et al.
Veröffentlicht: (2025)
von: Shutova, Alina, et al.
Veröffentlicht: (2025)
CAGE: Curvature-Aware Gradient Estimation For Accurate Quantization-Aware Training
von: Tabesh, Soroush, et al.
Veröffentlicht: (2025)
von: Tabesh, Soroush, et al.
Veröffentlicht: (2025)
LLMQ: Efficient Lower-Precision Pretraining for Consumer GPUs
von: Schultheis, Erik, et al.
Veröffentlicht: (2025)
von: Schultheis, Erik, et al.
Veröffentlicht: (2025)
A Hardware-Aware, Per-Layer Methodology for Post-Training Quantization of Large Language Models
von: Killian, Earl
Veröffentlicht: (2026)
von: Killian, Earl
Veröffentlicht: (2026)
Hogwild! Inference: Parallel LLM Generation via Concurrent Attention
von: Rodionov, Gleb, et al.
Veröffentlicht: (2025)
von: Rodionov, Gleb, et al.
Veröffentlicht: (2025)
Beyond Outliers: A Study of Optimizers Under Quantization
von: Vlassis, Georgios, et al.
Veröffentlicht: (2025)
von: Vlassis, Georgios, et al.
Veröffentlicht: (2025)
MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models
von: Frantar, Elias, et al.
Veröffentlicht: (2024)
von: Frantar, Elias, et al.
Veröffentlicht: (2024)
Statistically-Lossless Quantization of Large Language Models
von: Helcig, Michael, et al.
Veröffentlicht: (2026)
von: Helcig, Michael, et al.
Veröffentlicht: (2026)
The Geometry of LLM Quantization: GPTQ as Babai's Nearest Plane Algorithm
von: Chen, Jiale, et al.
Veröffentlicht: (2025)
von: Chen, Jiale, et al.
Veröffentlicht: (2025)
Quartet: Native FP4 Training Can Be Optimal for Large Language Models
von: Castro, Roberto L., et al.
Veröffentlicht: (2025)
von: Castro, Roberto L., et al.
Veröffentlicht: (2025)
FFT-based Dynamic Subspace Selection for Low-Rank Adaptive Optimization of Large Language Models
von: Modoranu, Ionut-Vlad, et al.
Veröffentlicht: (2025)
von: Modoranu, Ionut-Vlad, et al.
Veröffentlicht: (2025)
QuEST: Stable Training of LLMs with 1-Bit Weights and Activations
von: Panferov, Andrei, et al.
Veröffentlicht: (2025)
von: Panferov, Andrei, et al.
Veröffentlicht: (2025)
DASH: Faster Shampoo via Batched Block Preconditioning and Efficient Inverse-Root Solvers
von: Modoranu, Ionut-Vlad, et al.
Veröffentlicht: (2026)
von: Modoranu, Ionut-Vlad, et al.
Veröffentlicht: (2026)
HALO: Hadamard-Assisted Lower-Precision Optimization for LLMs
von: Ashkboos, Saleh, et al.
Veröffentlicht: (2025)
von: Ashkboos, Saleh, et al.
Veröffentlicht: (2025)
BPDQ: Bit-Plane Decomposition Quantization on a Variable Grid for Large Language Models
von: Chen, Junyu, et al.
Veröffentlicht: (2026)
von: Chen, Junyu, et al.
Veröffentlicht: (2026)
MatGPTQ: Accurate and Efficient Post-Training Matryoshka Quantization
von: Kleinegger, Maximilian, et al.
Veröffentlicht: (2026)
von: Kleinegger, Maximilian, et al.
Veröffentlicht: (2026)
Taming Unbalanced Training Workloads in Deep Learning with Partial Collective Operations
von: Li, Shigang, et al.
Veröffentlicht: (2019)
von: Li, Shigang, et al.
Veröffentlicht: (2019)
LeanQuant: Accurate and Scalable Large Language Model Quantization with Loss-error-aware Grid
von: Zhang, Tianyi, et al.
Veröffentlicht: (2024)
von: Zhang, Tianyi, et al.
Veröffentlicht: (2024)
Neural Optimal Transport with General Cost Functionals
von: Asadulaev, Arip, et al.
Veröffentlicht: (2022)
von: Asadulaev, Arip, et al.
Veröffentlicht: (2022)
Correlated Quantization for Faster Nonconvex Distributed Optimization
von: Panferov, Andrei, et al.
Veröffentlicht: (2024)
von: Panferov, Andrei, et al.
Veröffentlicht: (2024)
Mitigating the Impact of Outlier Channels for Language Model Quantization with Activation Regularization
von: Nrusimha, Aniruddha, et al.
Veröffentlicht: (2024)
von: Nrusimha, Aniruddha, et al.
Veröffentlicht: (2024)
Model Compression with Exact Budget Constraints via Riemannian Manifolds
von: Helcig, Michael, et al.
Veröffentlicht: (2026)
von: Helcig, Michael, et al.
Veröffentlicht: (2026)
Active Model Selection for Large Language Models
von: Durmazkeser, Yavuz, et al.
Veröffentlicht: (2025)
von: Durmazkeser, Yavuz, et al.
Veröffentlicht: (2025)
EfQAT: An Efficient Framework for Quantization-Aware Training
von: Ashkboos, Saleh, et al.
Veröffentlicht: (2024)
von: Ashkboos, Saleh, et al.
Veröffentlicht: (2024)
Rethinking Optimal Transport in Offline Reinforcement Learning
von: Asadulaev, Arip, et al.
Veröffentlicht: (2024)
von: Asadulaev, Arip, et al.
Veröffentlicht: (2024)
Mathador-LM: A Dynamic Benchmark for Mathematical Reasoning on Large Language Models
von: Kurtic, Eldar, et al.
Veröffentlicht: (2024)
von: Kurtic, Eldar, et al.
Veröffentlicht: (2024)
Large Language Model Selection with Limited Annotations
von: Durmazkeser, Yavuz, et al.
Veröffentlicht: (2026)
von: Durmazkeser, Yavuz, et al.
Veröffentlicht: (2026)
Epidemiology of Large Language Models: A Benchmark for Observational Distribution Knowledge
von: Plecko, Drago, et al.
Veröffentlicht: (2025)
von: Plecko, Drago, et al.
Veröffentlicht: (2025)
DarwinLM: Evolutionary Structured Pruning of Large Language Models
von: Tang, Shengkun, et al.
Veröffentlicht: (2025)
von: Tang, Shengkun, et al.
Veröffentlicht: (2025)
Vector Quantization in the Brain: Grid-like Codes in World Models
von: Peng, Xiangyuan, et al.
Veröffentlicht: (2025)
von: Peng, Xiangyuan, et al.
Veröffentlicht: (2025)
Accurate Compression of Text-to-Image Diffusion Models via Vector Quantization
von: Egiazarian, Vage, et al.
Veröffentlicht: (2024)
von: Egiazarian, Vage, et al.
Veröffentlicht: (2024)
RL2Grid: Benchmarking Reinforcement Learning in Power Grid Operations
von: Marchesini, Enrico, et al.
Veröffentlicht: (2025)
von: Marchesini, Enrico, et al.
Veröffentlicht: (2025)
QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs
von: Ashkboos, Saleh, et al.
Veröffentlicht: (2024)
von: Ashkboos, Saleh, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Extreme Compression of Large Language Models via Additive Quantization
von: Egiazarian, Vage, et al.
Veröffentlicht: (2024) -
WUSH: Near-Optimal Adaptive Transforms for LLM Quantization
von: Chen, Jiale, et al.
Veröffentlicht: (2025) -
Quartet II: Accurate LLM Pre-Training in NVFP4 by Improved Unbiased Gradient Estimation
von: Panferov, Andrei, et al.
Veröffentlicht: (2026) -
Bridging the Gap Between Promise and Performance for Microscaling FP4 Quantization
von: Egiazarian, Vage, et al.
Veröffentlicht: (2025) -
Unified Scaling Laws for Compressed Representations
von: Panferov, Andrei, et al.
Veröffentlicht: (2025)