Signal Processing Meets SGD: From Momentum to Filter

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yao, Zhipeng, Yu, Rui, Chang, Guisong, Li, Ying, Zhang, Yu, Li, Dazhou
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929747939819520
author Yao, Zhipeng
Yu, Rui
Chang, Guisong
Li, Ying
Zhang, Yu
Li, Dazhou
author_facet Yao, Zhipeng
Yu, Rui
Chang, Guisong
Li, Ying
Zhang, Yu
Li, Dazhou
contents In deep learning, stochastic gradient descent (SGD) and its momentum-based variants are widely used for optimization. However, the internal dynamics of these methods remain underexplored. In this paper, we analyze gradient behavior through a signal processing lens, isolating key factors that influence gradient updates and revealing a critical limitation: momentum techniques lack the flexibility to adequately balance bias and variance components in gradients, resulting in gradient estimation inaccuracies. To address this issue, we introduce a novel method SGDF (SGD with Filter) based on Wiener Filter principles, which derives an optimal time-varying gain to refine gradient updates by minimizing the mean square error in gradient estimation. This method yields an optimal first-order gradient estimate, effectively balancing noise reduction and signal preservation. Furthermore, our approach could extend to adaptive optimizers, enhancing their generalization potential. Empirical results show that SGDF achieves superior convergence and generalization compared to traditional momentum methods, and performs competitively with state-of-the-art optimizers.
format Preprint
id arxiv_https___arxiv_org_abs_2311_02818
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Signal Processing Meets SGD: From Momentum to Filter
Yao, Zhipeng
Yu, Rui
Chang, Guisong
Li, Ying
Zhang, Yu
Li, Dazhou
Machine Learning
Signal Processing
In deep learning, stochastic gradient descent (SGD) and its momentum-based variants are widely used for optimization. However, the internal dynamics of these methods remain underexplored. In this paper, we analyze gradient behavior through a signal processing lens, isolating key factors that influence gradient updates and revealing a critical limitation: momentum techniques lack the flexibility to adequately balance bias and variance components in gradients, resulting in gradient estimation inaccuracies. To address this issue, we introduce a novel method SGDF (SGD with Filter) based on Wiener Filter principles, which derives an optimal time-varying gain to refine gradient updates by minimizing the mean square error in gradient estimation. This method yields an optimal first-order gradient estimate, effectively balancing noise reduction and signal preservation. Furthermore, our approach could extend to adaptive optimizers, enhancing their generalization potential. Empirical results show that SGDF achieves superior convergence and generalization compared to traditional momentum methods, and performs competitively with state-of-the-art optimizers.
title Signal Processing Meets SGD: From Momentum to Filter
topic Machine Learning
Signal Processing
url https://arxiv.org/abs/2311.02818