HOT: Hadamard-based Optimized Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Seonggon, Shin, Juncheol, Woo, Seung-taek, Park, Eunhyeok
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913762334736384
author Kim, Seonggon
Shin, Juncheol
Woo, Seung-taek
Park, Eunhyeok
author_facet Kim, Seonggon
Shin, Juncheol
Woo, Seung-taek
Park, Eunhyeok
contents It has become increasingly important to optimize backpropagation to reduce memory usage and computational overhead. Achieving this goal is highly challenging, as multiple objectives must be considered jointly while maintaining training quality. In this paper, we focus on matrix multiplication, which accounts for the largest portion of training costs, and analyze its backpropagation in detail to identify lightweight techniques that offer the best benefits. Based on this analysis, we introduce a novel method, Hadamard-based Optimized Training (HOT). In this approach, we apply Hadamard-based optimizations, such as Hadamard quantization and Hadamard low-rank approximation, selectively and with awareness of the suitability of each optimization for different backward paths. Additionally, we introduce two enhancements: activation buffer compression and layer-wise quantizer selection. Our extensive analysis shows that HOT achieves up to 75% memory savings and a 2.6 times acceleration on real GPUs, with negligible accuracy loss compared to FP32 precision.
format Preprint
id arxiv_https___arxiv_org_abs_2503_21261
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HOT: Hadamard-based Optimized Training
Kim, Seonggon
Shin, Juncheol
Woo, Seung-taek
Park, Eunhyeok
Machine Learning
It has become increasingly important to optimize backpropagation to reduce memory usage and computational overhead. Achieving this goal is highly challenging, as multiple objectives must be considered jointly while maintaining training quality. In this paper, we focus on matrix multiplication, which accounts for the largest portion of training costs, and analyze its backpropagation in detail to identify lightweight techniques that offer the best benefits. Based on this analysis, we introduce a novel method, Hadamard-based Optimized Training (HOT). In this approach, we apply Hadamard-based optimizations, such as Hadamard quantization and Hadamard low-rank approximation, selectively and with awareness of the suitability of each optimization for different backward paths. Additionally, we introduce two enhancements: activation buffer compression and layer-wise quantizer selection. Our extensive analysis shows that HOT achieves up to 75% memory savings and a 2.6 times acceleration on real GPUs, with negligible accuracy loss compared to FP32 precision.
title HOT: Hadamard-based Optimized Training
topic Machine Learning
url https://arxiv.org/abs/2503.21261