Dynamic Decoupling of Placid Terminal Attractor-based Gradient Descent Algorithm

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Jinwei, Gori, Marco, Betti, Alessandro, Melacci, Stefano, Zhang, Hongtao, Liu, Jiedong, Hei, Xinhong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910597919014912
author Zhao, Jinwei
Gori, Marco
Betti, Alessandro
Melacci, Stefano
Zhang, Hongtao
Liu, Jiedong
Hei, Xinhong
author_facet Zhao, Jinwei
Gori, Marco
Betti, Alessandro
Melacci, Stefano
Zhang, Hongtao
Liu, Jiedong
Hei, Xinhong
contents Gradient descent (GD) and stochastic gradient descent (SGD) have been widely used in a large number of application domains. Therefore, understanding the dynamics of GD and improving its convergence speed is still of great importance. This paper carefully analyzes the dynamics of GD based on the terminal attractor at different stages of its gradient flow. On the basis of the terminal sliding mode theory and the terminal attractor theory, four adaptive learning rates are designed. Their performances are investigated in light of a detailed theoretical investigation, and the running times of the learning procedures are evaluated and compared. The total times of their learning processes are also studied in detail. To evaluate their effectiveness, various simulation results are investigated on a function approximation problem and an image classification problem.
format Preprint
id arxiv_https___arxiv_org_abs_2409_06542
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Dynamic Decoupling of Placid Terminal Attractor-based Gradient Descent Algorithm
Zhao, Jinwei
Gori, Marco
Betti, Alessandro
Melacci, Stefano
Zhang, Hongtao
Liu, Jiedong
Hei, Xinhong
Machine Learning
Gradient descent (GD) and stochastic gradient descent (SGD) have been widely used in a large number of application domains. Therefore, understanding the dynamics of GD and improving its convergence speed is still of great importance. This paper carefully analyzes the dynamics of GD based on the terminal attractor at different stages of its gradient flow. On the basis of the terminal sliding mode theory and the terminal attractor theory, four adaptive learning rates are designed. Their performances are investigated in light of a detailed theoretical investigation, and the running times of the learning procedures are evaluated and compared. The total times of their learning processes are also studied in detail. To evaluate their effectiveness, various simulation results are investigated on a function approximation problem and an image classification problem.
title Dynamic Decoupling of Placid Terminal Attractor-based Gradient Descent Algorithm
topic Machine Learning
url https://arxiv.org/abs/2409.06542