LOTION: Smoothing the Optimization Landscape for Quantized Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kwun, Mujin, Morwani, Depen, Su, Chloe Huangyuan, Gil, Stephanie, Anand, Nikhil, Kakade, Sham
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918157867810816
author Kwun, Mujin
Morwani, Depen
Su, Chloe Huangyuan
Gil, Stephanie
Anand, Nikhil
Kakade, Sham
author_facet Kwun, Mujin
Morwani, Depen
Su, Chloe Huangyuan
Gil, Stephanie
Anand, Nikhil
Kakade, Sham
contents Optimizing neural networks for quantized objectives is fundamentally challenging because the quantizer is piece-wise constant, yielding zero gradients everywhere except at quantization thresholds where the derivative is undefined. Most existing methods deal with this issue by relaxing gradient computations with techniques like Straight Through Estimators (STE) and do not provide any guarantees of convergence. In this work, taking inspiration from Nesterov smoothing, we approximate the quantized loss surface with a continuous loss surface. In particular, we introduce LOTION, \textbf{L}ow-precision \textbf{O}ptimization via s\textbf{T}ochastic-no\textbf{I}se sm\textbf{O}othi\textbf{N}g, a principled smoothing framework that replaces the raw quantized loss with its expectation under unbiased randomized-rounding noise. In this framework, standard optimizers are guaranteed to converge to a local minimum of the loss surface. Moreover, when using noise derived from stochastic rounding, we show that the global minima of the original quantized loss are preserved. We empirically demonstrate that this method outperforms standard QAT on synthetic testbeds and on 150M- and 300M- parameter language models.
format Preprint
id arxiv_https___arxiv_org_abs_2510_08757
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LOTION: Smoothing the Optimization Landscape for Quantized Training
Kwun, Mujin
Morwani, Depen
Su, Chloe Huangyuan
Gil, Stephanie
Anand, Nikhil
Kakade, Sham
Machine Learning
Hardware Architecture
Optimizing neural networks for quantized objectives is fundamentally challenging because the quantizer is piece-wise constant, yielding zero gradients everywhere except at quantization thresholds where the derivative is undefined. Most existing methods deal with this issue by relaxing gradient computations with techniques like Straight Through Estimators (STE) and do not provide any guarantees of convergence. In this work, taking inspiration from Nesterov smoothing, we approximate the quantized loss surface with a continuous loss surface. In particular, we introduce LOTION, \textbf{L}ow-precision \textbf{O}ptimization via s\textbf{T}ochastic-no\textbf{I}se sm\textbf{O}othi\textbf{N}g, a principled smoothing framework that replaces the raw quantized loss with its expectation under unbiased randomized-rounding noise. In this framework, standard optimizers are guaranteed to converge to a local minimum of the loss surface. Moreover, when using noise derived from stochastic rounding, we show that the global minima of the original quantized loss are preserved. We empirically demonstrate that this method outperforms standard QAT on synthetic testbeds and on 150M- and 300M- parameter language models.
title LOTION: Smoothing the Optimization Landscape for Quantized Training
topic Machine Learning
Hardware Architecture
url https://arxiv.org/abs/2510.08757