Layer-wise Quantization for Quantized Optimistic Dual Averaging
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916746293673984 |
|---|---|
| author | Nguyen, Anh Duc Markov, Ilia Wu, Frank Zhengqing Ramezani-Kebrya, Ali Antonakopoulos, Kimon Alistarh, Dan Cevher, Volkan |
| author_facet | Nguyen, Anh Duc Markov, Ilia Wu, Frank Zhengqing Ramezani-Kebrya, Ali Antonakopoulos, Kimon Alistarh, Dan Cevher, Volkan |
| contents | Modern deep neural networks exhibit heterogeneity across numerous layers of various types such as residuals, multi-head attention, etc., due to varying structures (dimensions, activation functions, etc.), distinct representation characteristics, which impact predictions. We develop a general layer-wise quantization framework with tight variance and code-length bounds, adapting to the heterogeneities over the course of training. We then apply a new layer-wise quantization technique within distributed variational inequalities (VIs), proposing a novel Quantized Optimistic Dual Averaging (QODA) algorithm with adaptive learning rates, which achieves competitive convergence rates for monotone VIs. We empirically show that QODA achieves up to a $150\%$ speedup over the baselines in end-to-end training time for training Wasserstein GAN on $12+$ GPUs. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_14371 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Layer-wise Quantization for Quantized Optimistic Dual Averaging Nguyen, Anh Duc Markov, Ilia Wu, Frank Zhengqing Ramezani-Kebrya, Ali Antonakopoulos, Kimon Alistarh, Dan Cevher, Volkan Machine Learning Optimization and Control Modern deep neural networks exhibit heterogeneity across numerous layers of various types such as residuals, multi-head attention, etc., due to varying structures (dimensions, activation functions, etc.), distinct representation characteristics, which impact predictions. We develop a general layer-wise quantization framework with tight variance and code-length bounds, adapting to the heterogeneities over the course of training. We then apply a new layer-wise quantization technique within distributed variational inequalities (VIs), proposing a novel Quantized Optimistic Dual Averaging (QODA) algorithm with adaptive learning rates, which achieves competitive convergence rates for monotone VIs. We empirically show that QODA achieves up to a $150\%$ speedup over the baselines in end-to-end training time for training Wasserstein GAN on $12+$ GPUs. |
| title | Layer-wise Quantization for Quantized Optimistic Dual Averaging |
| topic | Machine Learning Optimization and Control |
| url | https://arxiv.org/abs/2505.14371 |