Adaptive Momentum and Nonlinear Damping for Neural Network Training
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918316490096640 |
|---|---|
| author | Karoni, Aikaterini Rajpal, Rajit Leimkuhler, Benedict Stoltz, Gabriel |
| author_facet | Karoni, Aikaterini Rajpal, Rajit Leimkuhler, Benedict Stoltz, Gabriel |
| contents | We propose a continuous-time scheme for large-scale optimization that introduces individual, adaptive momentum coefficients regulated by the kinetic energy of each model parameter. This approach automatically adjusts to local landscape curvature to maintain stability without sacrificing convergence speed. We demonstrate that our adaptive friction can be related to cubic damping, a suppression mechanism from structural dynamics. Furthermore, we introduce two specific optimization schemes by augmenting the continuous dynamics of mSGD and Adam with a cubic damping term. Empirically, our methods demonstrate robustness and match or outperform Adam on training ViT, BERT, and GPT2 tasks where mSGD typically struggles. We further provide theoretical results establishing the exponential convergence of the proposed schemes. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_00334 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Adaptive Momentum and Nonlinear Damping for Neural Network Training Karoni, Aikaterini Rajpal, Rajit Leimkuhler, Benedict Stoltz, Gabriel Machine Learning Optimization and Control We propose a continuous-time scheme for large-scale optimization that introduces individual, adaptive momentum coefficients regulated by the kinetic energy of each model parameter. This approach automatically adjusts to local landscape curvature to maintain stability without sacrificing convergence speed. We demonstrate that our adaptive friction can be related to cubic damping, a suppression mechanism from structural dynamics. Furthermore, we introduce two specific optimization schemes by augmenting the continuous dynamics of mSGD and Adam with a cubic damping term. Empirically, our methods demonstrate robustness and match or outperform Adam on training ViT, BERT, and GPT2 tasks where mSGD typically struggles. We further provide theoretical results establishing the exponential convergence of the proposed schemes. |
| title | Adaptive Momentum and Nonlinear Damping for Neural Network Training |
| topic | Machine Learning Optimization and Control |
| url | https://arxiv.org/abs/2602.00334 |