Does Weight Decay Enhance Training Stability?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Saether, Marius, Kolic, Amir, Poggio, Tomaso, Beneventano, Pierfrancesco |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
How Neural Networks Learn the Support is an Implicit Regularization Effect of SGD
von: Beneventano, Pierfrancesco, et al.
Veröffentlicht: (2024)
von: Beneventano, Pierfrancesco, et al.
Veröffentlicht: (2024)
Momentum Further Constrains Sharpness at the Edge of Stochastic Stability
von: Andreyev, Arseniy, et al.
Veröffentlicht: (2026)
von: Andreyev, Arseniy, et al.
Veröffentlicht: (2026)
Too Sharp, Too Sure: When Calibration Follows Curvature
von: Morosini, Alessandro, et al.
Veröffentlicht: (2026)
von: Morosini, Alessandro, et al.
Veröffentlicht: (2026)
Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias
von: Das, Mohua, et al.
Veröffentlicht: (2026)
von: Das, Mohua, et al.
Veröffentlicht: (2026)
Edge of Stochastic Stability: Revisiting the Edge of Stability for SGD
von: Andreyev, Arseniy, et al.
Veröffentlicht: (2024)
von: Andreyev, Arseniy, et al.
Veröffentlicht: (2024)
On the Trajectories of SGD Without Replacement
von: Beneventano, Pierfrancesco
Veröffentlicht: (2023)
von: Beneventano, Pierfrancesco
Veröffentlicht: (2023)
Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks
von: Beneventano, Pierfrancesco, et al.
Veröffentlicht: (2025)
von: Beneventano, Pierfrancesco, et al.
Veröffentlicht: (2025)
Does SGD Seek Flatness or Sharpness? An Exactly Solvable Model
von: Xu, Yizhou, et al.
Veröffentlicht: (2026)
von: Xu, Yizhou, et al.
Veröffentlicht: (2026)
Iterative regularization in classification via hinge loss diagonal descent
von: Apidopoulos, Vassilis, et al.
Veröffentlicht: (2022)
von: Apidopoulos, Vassilis, et al.
Veröffentlicht: (2022)
pAI/MSc: ML Theory Research with Humans on the Loop
von: Abdelmoneum, Mahmoud, et al.
Veröffentlicht: (2026)
von: Abdelmoneum, Mahmoud, et al.
Veröffentlicht: (2026)
Cautious Weight Decay
von: Chen, Lizhang, et al.
Veröffentlicht: (2025)
von: Chen, Lizhang, et al.
Veröffentlicht: (2025)
Wide Neural Networks Trained with Weight Decay Provably Exhibit Neural Collapse
von: Jacot, Arthur, et al.
Veröffentlicht: (2024)
von: Jacot, Arthur, et al.
Veröffentlicht: (2024)
Solving Inverse Problems with Deep Linear Neural Networks: Global Convergence Guarantees for Gradient Descent with Weight Decay
von: Laus, Hannah, et al.
Veröffentlicht: (2025)
von: Laus, Hannah, et al.
Veröffentlicht: (2025)
Enhancing Stability of Physics-Informed Neural Network Training Through Saddle-Point Reformulation
von: Bylinkin, Dmitry, et al.
Veröffentlicht: (2025)
von: Bylinkin, Dmitry, et al.
Veröffentlicht: (2025)
Same Error, Different Function: The Optimizer as an Implicit Prior in Financial Time Series
von: Cortesi, Federico Vittorio, et al.
Veröffentlicht: (2026)
von: Cortesi, Federico Vittorio, et al.
Veröffentlicht: (2026)
Controlling the Flow: Stability and Convergence for Stochastic Gradient Descent with Decaying Regularization
von: Kassing, Sebastian, et al.
Veröffentlicht: (2025)
von: Kassing, Sebastian, et al.
Veröffentlicht: (2025)
GNMR: Runtime Stability Control for Low-Precision Large Language Model Training
von: Kong, Boao, et al.
Veröffentlicht: (2026)
von: Kong, Boao, et al.
Veröffentlicht: (2026)
An Adaptive and Stability-Promoting Layerwise Training Approach for Sparse Deep Neural Network Architecture
von: Krishnanunni, C G, et al.
Veröffentlicht: (2022)
von: Krishnanunni, C G, et al.
Veröffentlicht: (2022)
Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size
von: Islamov, Rustem, et al.
Veröffentlicht: (2025)
von: Islamov, Rustem, et al.
Veröffentlicht: (2025)
Enhancing Privacy in Federated Learning through Local Training
von: Bastianello, Nicola, et al.
Veröffentlicht: (2024)
von: Bastianello, Nicola, et al.
Veröffentlicht: (2024)
Decoupled Weight Decay for Any $p$ Norm
von: Outmezguine, Nadav Joseph, et al.
Veröffentlicht: (2024)
von: Outmezguine, Nadav Joseph, et al.
Veröffentlicht: (2024)
Tight Long-Term Tail Decay of (Clipped) SGD in Non-Convex Optimization
von: Armacki, Aleksandar, et al.
Veröffentlicht: (2026)
von: Armacki, Aleksandar, et al.
Veröffentlicht: (2026)
Convergence of Sharpness-Aware Minimization Algorithms using Increasing Batch Size and Decaying Learning Rate
von: Harada, Hinata, et al.
Veröffentlicht: (2024)
von: Harada, Hinata, et al.
Veröffentlicht: (2024)
Muon Does Not Converge on Convex Lipschitz Functions
von: Parshakova, Tetiana, et al.
Veröffentlicht: (2026)
von: Parshakova, Tetiana, et al.
Veröffentlicht: (2026)
Does SGD really happen in tiny subspaces?
von: Song, Minhak, et al.
Veröffentlicht: (2024)
von: Song, Minhak, et al.
Veröffentlicht: (2024)
From Cursed to Competitive: Closing the ZO-FO Gap via Input-to-State Stability
von: Farzin, Amir Ali, et al.
Veröffentlicht: (2026)
von: Farzin, Amir Ali, et al.
Veröffentlicht: (2026)
Gauss-Newton Natural Gradient Descent for Shape Learning
von: King, James, et al.
Veröffentlicht: (2026)
von: King, James, et al.
Veröffentlicht: (2026)
An Inexact Weighted Proximal Trust-Region Method
von: Maia, Leandro Farias, et al.
Veröffentlicht: (2026)
von: Maia, Leandro Farias, et al.
Veröffentlicht: (2026)
Robust Implicit Regularization via Weight Normalization
von: Chou, Hung-Hsu, et al.
Veröffentlicht: (2023)
von: Chou, Hung-Hsu, et al.
Veröffentlicht: (2023)
A Unified Analysis for Finite Weight Averaging
von: Wang, Peng, et al.
Veröffentlicht: (2024)
von: Wang, Peng, et al.
Veröffentlicht: (2024)
Momentum Does Not Reduce Stochastic Noise in Stochastic Gradient Descent
von: Sato, Naoki, et al.
Veröffentlicht: (2024)
von: Sato, Naoki, et al.
Veröffentlicht: (2024)
LORENZA: Enhancing Generalization in Low-Rank Gradient LLM Training via Efficient Zeroth-Order Adaptive SAM
von: Refael, Yehonathan, et al.
Veröffentlicht: (2025)
von: Refael, Yehonathan, et al.
Veröffentlicht: (2025)
Limits of Convergence-Rate Control for Open-Weight Safety
von: Rosati, Domenic, et al.
Veröffentlicht: (2026)
von: Rosati, Domenic, et al.
Veröffentlicht: (2026)
WeightLoRA: Keep Only Necessary Adapters
von: Veprikov, Andrey, et al.
Veröffentlicht: (2025)
von: Veprikov, Andrey, et al.
Veröffentlicht: (2025)
Implicit Differentiation for Hyperparameter Tuning the Weighted Graphical Lasso
von: Pouliquen, Can, et al.
Veröffentlicht: (2023)
von: Pouliquen, Can, et al.
Veröffentlicht: (2023)
Minimisation of Polyak-Łojasewicz Functions Using Random Zeroth-Order Oracles
von: Farzin, Amir Ali, et al.
Veröffentlicht: (2024)
von: Farzin, Amir Ali, et al.
Veröffentlicht: (2024)
Offline Policy Learning with Weight Clipping and Heaviside Composite Optimization
von: Liu, Jingren, et al.
Veröffentlicht: (2026)
von: Liu, Jingren, et al.
Veröffentlicht: (2026)
Efficient Alternating Minimization with Applications to Weighted Low Rank Approximation
von: Song, Zhao, et al.
Veröffentlicht: (2023)
von: Song, Zhao, et al.
Veröffentlicht: (2023)
A Simplified Analysis of SGD for Linear Regression with Weight Averaging
von: Meterez, Alexandru, et al.
Veröffentlicht: (2025)
von: Meterez, Alexandru, et al.
Veröffentlicht: (2025)
On the Stability Connection Between Discrete-Time Algorithms and Their Resolution ODEs: Applications to Min-Max Optimisation
von: Farzin, Amir Ali, et al.
Veröffentlicht: (2026)
von: Farzin, Amir Ali, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
How Neural Networks Learn the Support is an Implicit Regularization Effect of SGD
von: Beneventano, Pierfrancesco, et al.
Veröffentlicht: (2024) -
Momentum Further Constrains Sharpness at the Edge of Stochastic Stability
von: Andreyev, Arseniy, et al.
Veröffentlicht: (2026) -
Too Sharp, Too Sure: When Calibration Follows Curvature
von: Morosini, Alessandro, et al.
Veröffentlicht: (2026) -
Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias
von: Das, Mohua, et al.
Veröffentlicht: (2026) -
Edge of Stochastic Stability: Revisiting the Edge of Stability for SGD
von: Andreyev, Arseniy, et al.
Veröffentlicht: (2024)