Adaptive Momentum and Nonlinear Damping for Neural Network Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Karoni, Aikaterini, Rajpal, Rajit, Leimkuhler, Benedict, Stoltz, Gabriel
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918316490096640
author Karoni, Aikaterini
Rajpal, Rajit
Leimkuhler, Benedict
Stoltz, Gabriel
author_facet Karoni, Aikaterini
Rajpal, Rajit
Leimkuhler, Benedict
Stoltz, Gabriel
contents We propose a continuous-time scheme for large-scale optimization that introduces individual, adaptive momentum coefficients regulated by the kinetic energy of each model parameter. This approach automatically adjusts to local landscape curvature to maintain stability without sacrificing convergence speed. We demonstrate that our adaptive friction can be related to cubic damping, a suppression mechanism from structural dynamics. Furthermore, we introduce two specific optimization schemes by augmenting the continuous dynamics of mSGD and Adam with a cubic damping term. Empirically, our methods demonstrate robustness and match or outperform Adam on training ViT, BERT, and GPT2 tasks where mSGD typically struggles. We further provide theoretical results establishing the exponential convergence of the proposed schemes.
format Preprint
id arxiv_https___arxiv_org_abs_2602_00334
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Adaptive Momentum and Nonlinear Damping for Neural Network Training
Karoni, Aikaterini
Rajpal, Rajit
Leimkuhler, Benedict
Stoltz, Gabriel
Machine Learning
Optimization and Control
We propose a continuous-time scheme for large-scale optimization that introduces individual, adaptive momentum coefficients regulated by the kinetic energy of each model parameter. This approach automatically adjusts to local landscape curvature to maintain stability without sacrificing convergence speed. We demonstrate that our adaptive friction can be related to cubic damping, a suppression mechanism from structural dynamics. Furthermore, we introduce two specific optimization schemes by augmenting the continuous dynamics of mSGD and Adam with a cubic damping term. Empirically, our methods demonstrate robustness and match or outperform Adam on training ViT, BERT, and GPT2 tasks where mSGD typically struggles. We further provide theoretical results establishing the exponential convergence of the proposed schemes.
title Adaptive Momentum and Nonlinear Damping for Neural Network Training
topic Machine Learning
Optimization and Control
url https://arxiv.org/abs/2602.00334