Layer-wise Quantization for Quantized Optimistic Dual Averaging

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nguyen, Anh Duc, Markov, Ilia, Wu, Frank Zhengqing, Ramezani-Kebrya, Ali, Antonakopoulos, Kimon, Alistarh, Dan, Cevher, Volkan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916746293673984
author Nguyen, Anh Duc
Markov, Ilia
Wu, Frank Zhengqing
Ramezani-Kebrya, Ali
Antonakopoulos, Kimon
Alistarh, Dan
Cevher, Volkan
author_facet Nguyen, Anh Duc
Markov, Ilia
Wu, Frank Zhengqing
Ramezani-Kebrya, Ali
Antonakopoulos, Kimon
Alistarh, Dan
Cevher, Volkan
contents Modern deep neural networks exhibit heterogeneity across numerous layers of various types such as residuals, multi-head attention, etc., due to varying structures (dimensions, activation functions, etc.), distinct representation characteristics, which impact predictions. We develop a general layer-wise quantization framework with tight variance and code-length bounds, adapting to the heterogeneities over the course of training. We then apply a new layer-wise quantization technique within distributed variational inequalities (VIs), proposing a novel Quantized Optimistic Dual Averaging (QODA) algorithm with adaptive learning rates, which achieves competitive convergence rates for monotone VIs. We empirically show that QODA achieves up to a $150\%$ speedup over the baselines in end-to-end training time for training Wasserstein GAN on $12+$ GPUs.
format Preprint
id arxiv_https___arxiv_org_abs_2505_14371
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Layer-wise Quantization for Quantized Optimistic Dual Averaging
Nguyen, Anh Duc
Markov, Ilia
Wu, Frank Zhengqing
Ramezani-Kebrya, Ali
Antonakopoulos, Kimon
Alistarh, Dan
Cevher, Volkan
Machine Learning
Optimization and Control
Modern deep neural networks exhibit heterogeneity across numerous layers of various types such as residuals, multi-head attention, etc., due to varying structures (dimensions, activation functions, etc.), distinct representation characteristics, which impact predictions. We develop a general layer-wise quantization framework with tight variance and code-length bounds, adapting to the heterogeneities over the course of training. We then apply a new layer-wise quantization technique within distributed variational inequalities (VIs), proposing a novel Quantized Optimistic Dual Averaging (QODA) algorithm with adaptive learning rates, which achieves competitive convergence rates for monotone VIs. We empirically show that QODA achieves up to a $150\%$ speedup over the baselines in end-to-end training time for training Wasserstein GAN on $12+$ GPUs.
title Layer-wise Quantization for Quantized Optimistic Dual Averaging
topic Machine Learning
Optimization and Control
url https://arxiv.org/abs/2505.14371