HALO: Hadamard-Assisted Lower-Precision Optimization for LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Ashkboos, Saleh, Nikdan, Mahdi, Tabesh, Soroush, Castro, Roberto L., Hoefler, Torsten, Alistarh, Dan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Quartet: Native FP4 Training Can Be Optimal for Large Language Models
por: Castro, Roberto L., et al.
Publicado: (2025)
por: Castro, Roberto L., et al.
Publicado: (2025)
QuEST: Stable Training of LLMs with 1-Bit Weights and Activations
por: Panferov, Andrei, et al.
Publicado: (2025)
por: Panferov, Andrei, et al.
Publicado: (2025)
RoSA: Accurate Parameter-Efficient Fine-Tuning via Robust Adaptation
por: Nikdan, Mahdi, et al.
Publicado: (2024)
por: Nikdan, Mahdi, et al.
Publicado: (2024)
GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling
por: Dadgarnia, Alireza, et al.
Publicado: (2026)
por: Dadgarnia, Alireza, et al.
Publicado: (2026)
Beyond Outliers: A Study of Optimizers Under Quantization
por: Vlassis, Georgios, et al.
Publicado: (2025)
por: Vlassis, Georgios, et al.
Publicado: (2025)
ECO: Quantized Training without Full-Precision Master Weights
por: Nikdan, Mahdi, et al.
Publicado: (2026)
por: Nikdan, Mahdi, et al.
Publicado: (2026)
MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models
por: Frantar, Elias, et al.
Publicado: (2024)
por: Frantar, Elias, et al.
Publicado: (2024)
QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs
por: Ashkboos, Saleh, et al.
Publicado: (2024)
por: Ashkboos, Saleh, et al.
Publicado: (2024)
Quartet II: Accurate LLM Pre-Training in NVFP4 by Improved Unbiased Gradient Estimation
por: Panferov, Andrei, et al.
Publicado: (2026)
por: Panferov, Andrei, et al.
Publicado: (2026)
EfQAT: An Efficient Framework for Quantization-Aware Training
por: Ashkboos, Saleh, et al.
Publicado: (2024)
por: Ashkboos, Saleh, et al.
Publicado: (2024)
Efficient Data Selection at Scale via Influence Distillation
por: Nikdan, Mahdi, et al.
Publicado: (2025)
por: Nikdan, Mahdi, et al.
Publicado: (2025)
Bridging the Gap Between Promise and Performance for Microscaling FP4 Quantization
por: Egiazarian, Vage, et al.
Publicado: (2025)
por: Egiazarian, Vage, et al.
Publicado: (2025)
WUSH: Near-Optimal Adaptive Transforms for LLM Quantization
por: Chen, Jiale, et al.
Publicado: (2025)
por: Chen, Jiale, et al.
Publicado: (2025)
CAGE: Curvature-Aware Gradient Estimation For Accurate Quantization-Aware Training
por: Tabesh, Soroush, et al.
Publicado: (2025)
por: Tabesh, Soroush, et al.
Publicado: (2025)
SliceGPT: Compress Large Language Models by Deleting Rows and Columns
por: Ashkboos, Saleh, et al.
Publicado: (2024)
por: Ashkboos, Saleh, et al.
Publicado: (2024)
LLMQ: Efficient Lower-Precision Pretraining for Consumer GPUs
por: Schultheis, Erik, et al.
Publicado: (2025)
por: Schultheis, Erik, et al.
Publicado: (2025)
Grid Games: The Power of Multiple Grids for Quantizing Large Language Models
por: Egiazarian, Vage, et al.
Publicado: (2026)
por: Egiazarian, Vage, et al.
Publicado: (2026)
The Geometry of LLM Quantization: GPTQ as Babai's Nearest Plane Algorithm
por: Chen, Jiale, et al.
Publicado: (2025)
por: Chen, Jiale, et al.
Publicado: (2025)
Behemoth: Benchmarking Unlearning in LLMs Using Fully Synthetic Data
por: Iofinova, Eugenia, et al.
Publicado: (2026)
por: Iofinova, Eugenia, et al.
Publicado: (2026)
Taming Unbalanced Training Workloads in Deep Learning with Partial Collective Operations
por: Li, Shigang, et al.
Publicado: (2019)
por: Li, Shigang, et al.
Publicado: (2019)
Can LLMs Separate Instructions From Data? And What Do We Even Mean By That?
por: Zverev, Egor, et al.
Publicado: (2024)
por: Zverev, Egor, et al.
Publicado: (2024)
Panza: Design and Analysis of a Fully-Local Personalized Text Writing Assistant
por: Nicolicioiu, Armand, et al.
Publicado: (2024)
por: Nicolicioiu, Armand, et al.
Publicado: (2024)
Breaking (Global) Barriers in Parallel Stochastic Optimization with Wait-Avoiding Group Averaging
por: Li, Shigang, et al.
Publicado: (2020)
por: Li, Shigang, et al.
Publicado: (2020)
Model Compression with Exact Budget Constraints via Riemannian Manifolds
por: Helcig, Michael, et al.
Publicado: (2026)
por: Helcig, Michael, et al.
Publicado: (2026)
Towards Robust Scaling Laws for Optimizers
por: Volkova, Alexandra, et al.
Publicado: (2026)
por: Volkova, Alexandra, et al.
Publicado: (2026)
EntryPrune: Neural Network Feature Selection using First Impressions
por: Zimmer, Felix, et al.
Publicado: (2024)
por: Zimmer, Felix, et al.
Publicado: (2024)
MatGPTQ: Accurate and Efficient Post-Training Matryoshka Quantization
por: Kleinegger, Maximilian, et al.
Publicado: (2026)
por: Kleinegger, Maximilian, et al.
Publicado: (2026)
Statistically-Lossless Quantization of Large Language Models
por: Helcig, Michael, et al.
Publicado: (2026)
por: Helcig, Michael, et al.
Publicado: (2026)
Powerset Convolutional Neural Networks
por: Wendler, Chris, et al.
Publicado: (2019)
por: Wendler, Chris, et al.
Publicado: (2019)
Hybrid Decentralized Optimization: Leveraging Both First- and Zeroth-Order Optimizers for Faster Convergence
por: Ansaripour, Matin, et al.
Publicado: (2022)
por: Ansaripour, Matin, et al.
Publicado: (2022)
LDAdam: Adaptive Optimization from Low-Dimensional Gradient Statistics
por: Robert, Thomas, et al.
Publicado: (2024)
por: Robert, Thomas, et al.
Publicado: (2024)
Simple Opinion Dynamics for No-Regret Learning
por: Lazarsfeld, John, et al.
Publicado: (2023)
por: Lazarsfeld, John, et al.
Publicado: (2023)
Computational Bottlenecks of Training Small-scale Large Language Models
por: Ashkboos, Saleh, et al.
Publicado: (2024)
por: Ashkboos, Saleh, et al.
Publicado: (2024)
Scaling Laws of Global Weather Models
por: Yu, Yuejiang, et al.
Publicado: (2026)
por: Yu, Yuejiang, et al.
Publicado: (2026)
Near-Optimal Sparse Allreduce for Distributed Deep Learning
por: Li, Shigang, et al.
Publicado: (2022)
por: Li, Shigang, et al.
Publicado: (2022)
Chimera: Efficiently Training Large-Scale Neural Networks with Bidirectional Pipelines
por: Li, Shigang, et al.
Publicado: (2021)
por: Li, Shigang, et al.
Publicado: (2021)
Apertus LLM Family Expansion via Distillation and Quantization
por: Panferov, Andrei, et al.
Publicado: (2026)
por: Panferov, Andrei, et al.
Publicado: (2026)
EvoPress: Accurate Dynamic Model Compression via Evolutionary Search
por: Sieberling, Oliver, et al.
Publicado: (2024)
por: Sieberling, Oliver, et al.
Publicado: (2024)
Communication-Efficient Federated Learning With Data and Client Heterogeneity
por: Zakerinia, Hossein, et al.
Publicado: (2022)
por: Zakerinia, Hossein, et al.
Publicado: (2022)
HOT: Hadamard-based Optimized Training
por: Kim, Seonggon, et al.
Publicado: (2025)
por: Kim, Seonggon, et al.
Publicado: (2025)
Ejemplares similares
-
Quartet: Native FP4 Training Can Be Optimal for Large Language Models
por: Castro, Roberto L., et al.
Publicado: (2025) -
QuEST: Stable Training of LLMs with 1-Bit Weights and Activations
por: Panferov, Andrei, et al.
Publicado: (2025) -
RoSA: Accurate Parameter-Efficient Fine-Tuning via Robust Adaptation
por: Nikdan, Mahdi, et al.
Publicado: (2024) -
GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling
por: Dadgarnia, Alireza, et al.
Publicado: (2026) -
Beyond Outliers: A Study of Optimizers Under Quantization
por: Vlassis, Georgios, et al.
Publicado: (2025)