How Neural Networks Learn the Support is an Implicit Regularization Effect of SGD
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Beneventano, Pierfrancesco, Pinto, Andrea, Poggio, Tomaso |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
On the Trajectories of SGD Without Replacement
von: Beneventano, Pierfrancesco
Veröffentlicht: (2023)
von: Beneventano, Pierfrancesco
Veröffentlicht: (2023)
Does Weight Decay Enhance Training Stability?
von: Saether, Marius, et al.
Veröffentlicht: (2026)
von: Saether, Marius, et al.
Veröffentlicht: (2026)
Too Sharp, Too Sure: When Calibration Follows Curvature
von: Morosini, Alessandro, et al.
Veröffentlicht: (2026)
von: Morosini, Alessandro, et al.
Veröffentlicht: (2026)
Edge of Stochastic Stability: Revisiting the Edge of Stability for SGD
von: Andreyev, Arseniy, et al.
Veröffentlicht: (2024)
von: Andreyev, Arseniy, et al.
Veröffentlicht: (2024)
Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias
von: Das, Mohua, et al.
Veröffentlicht: (2026)
von: Das, Mohua, et al.
Veröffentlicht: (2026)
Momentum Further Constrains Sharpness at the Edge of Stochastic Stability
von: Andreyev, Arseniy, et al.
Veröffentlicht: (2026)
von: Andreyev, Arseniy, et al.
Veröffentlicht: (2026)
Does SGD Seek Flatness or Sharpness? An Exactly Solvable Model
von: Xu, Yizhou, et al.
Veröffentlicht: (2026)
von: Xu, Yizhou, et al.
Veröffentlicht: (2026)
Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks
von: Beneventano, Pierfrancesco, et al.
Veröffentlicht: (2025)
von: Beneventano, Pierfrancesco, et al.
Veröffentlicht: (2025)
Iterative regularization in classification via hinge loss diagonal descent
von: Apidopoulos, Vassilis, et al.
Veröffentlicht: (2022)
von: Apidopoulos, Vassilis, et al.
Veröffentlicht: (2022)
pAI/MSc: ML Theory Research with Humans on the Loop
von: Abdelmoneum, Mahmoud, et al.
Veröffentlicht: (2026)
von: Abdelmoneum, Mahmoud, et al.
Veröffentlicht: (2026)
The Rich and the Simple: On the Implicit Bias of Adam and SGD
von: Vasudeva, Bhavya, et al.
Veröffentlicht: (2025)
von: Vasudeva, Bhavya, et al.
Veröffentlicht: (2025)
SGD with Partial Hessian for Deep Neural Networks Optimization
von: Sun, Ying, et al.
Veröffentlicht: (2024)
von: Sun, Ying, et al.
Veröffentlicht: (2024)
Same Error, Different Function: The Optimizer as an Implicit Prior in Financial Time Series
von: Cortesi, Federico Vittorio, et al.
Veröffentlicht: (2026)
von: Cortesi, Federico Vittorio, et al.
Veröffentlicht: (2026)
Convergence of SGD for Training Neural Networks with Sliced Wasserstein Losses
von: Tanguy, Eloi
Veröffentlicht: (2023)
von: Tanguy, Eloi
Veröffentlicht: (2023)
Global Convergence of SGD On Two Layer Neural Nets
von: Gopalani, Pulkit, et al.
Veröffentlicht: (2022)
von: Gopalani, Pulkit, et al.
Veröffentlicht: (2022)
Improving Generalization and Convergence by Enhancing Implicit Regularization
von: Wang, Mingze, et al.
Veröffentlicht: (2024)
von: Wang, Mingze, et al.
Veröffentlicht: (2024)
Robust Implicit Regularization via Weight Normalization
von: Chou, Hung-Hsu, et al.
Veröffentlicht: (2023)
von: Chou, Hung-Hsu, et al.
Veröffentlicht: (2023)
Global Convergence of SGD For Logistic Loss on Two Layer Neural Nets
von: Gopalani, Pulkit, et al.
Veröffentlicht: (2023)
von: Gopalani, Pulkit, et al.
Veröffentlicht: (2023)
How Does the ReLU Activation Affect the Implicit Bias of Gradient Descent on High-dimensional Neural Network Regression?
von: Lai, Kuo-Wei, et al.
Veröffentlicht: (2026)
von: Lai, Kuo-Wei, et al.
Veröffentlicht: (2026)
The Marginal Value of Momentum for Small Learning Rate SGD
von: Wang, Runzhe, et al.
Veröffentlicht: (2023)
von: Wang, Runzhe, et al.
Veröffentlicht: (2023)
SLowcal-SGD: Slow Query Points Improve Local-SGD for Stochastic Convex Optimization
von: Dahan, Tehila, et al.
Veröffentlicht: (2023)
von: Dahan, Tehila, et al.
Veröffentlicht: (2023)
Regularized Gauss-Newton for Optimizing Overparameterized Neural Networks
von: Adeoye, Adeyemi D., et al.
Veröffentlicht: (2024)
von: Adeoye, Adeyemi D., et al.
Veröffentlicht: (2024)
Understanding the Implicit Regularization of Gradient Descent in Over-parameterized Models
von: Ma, Jianhao, et al.
Veröffentlicht: (2025)
von: Ma, Jianhao, et al.
Veröffentlicht: (2025)
Making SGD Parameter-Free
von: Carmon, Yair, et al.
Veröffentlicht: (2022)
von: Carmon, Yair, et al.
Veröffentlicht: (2022)
Shadowheart SGD: Distributed Asynchronous SGD with Optimal Time Complexity Under Arbitrary Computation and Communication Heterogeneity
von: Tyurin, Alexander, et al.
Veröffentlicht: (2024)
von: Tyurin, Alexander, et al.
Veröffentlicht: (2024)
Implicit Regularization Makes Overparameterized Asymmetric Matrix Sensing Robust to Perturbations
von: Wind, Johan S.
Veröffentlicht: (2023)
von: Wind, Johan S.
Veröffentlicht: (2023)
Implicit Regularization in Perturbed Deep Matrix Factorization: Spectral Conditions and Stability
von: Wang, Jingzhe, et al.
Veröffentlicht: (2026)
von: Wang, Jingzhe, et al.
Veröffentlicht: (2026)
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees
von: Xie, Shengping, et al.
Veröffentlicht: (2025)
von: Xie, Shengping, et al.
Veröffentlicht: (2025)
Understanding Outer Optimizers in Local SGD: Learning Rates, Momentum, and Acceleration
von: Khaled, Ahmed, et al.
Veröffentlicht: (2025)
von: Khaled, Ahmed, et al.
Veröffentlicht: (2025)
Demystifying SGD with Doubly Stochastic Gradients
von: Kim, Kyurae, et al.
Veröffentlicht: (2024)
von: Kim, Kyurae, et al.
Veröffentlicht: (2024)
Dimension-adapted Momentum Outscales SGD
von: Ferbach, Damien, et al.
Veröffentlicht: (2025)
von: Ferbach, Damien, et al.
Veröffentlicht: (2025)
Heavy-Tail Phenomenon in Decentralized SGD
von: Gurbuzbalaban, Mert, et al.
Veröffentlicht: (2022)
von: Gurbuzbalaban, Mert, et al.
Veröffentlicht: (2022)
Convergence of Implicit Gradient Descent for Training Two-Layer Physics-Informed Neural Networks
von: Xu, Xianliang, et al.
Veröffentlicht: (2024)
von: Xu, Xianliang, et al.
Veröffentlicht: (2024)
Regularized Gradient Clipping Provably Trains Wide and Deep Neural Networks
von: Tucat, Matteo, et al.
Veröffentlicht: (2024)
von: Tucat, Matteo, et al.
Veröffentlicht: (2024)
Neural Network Training Techniques Regularize Optimization Trajectory: An Empirical Study
von: Chen, Cheng, et al.
Veröffentlicht: (2020)
von: Chen, Cheng, et al.
Veröffentlicht: (2020)
SGD with memory: fundamental properties and stochastic acceleration
von: Yarotsky, Dmitry, et al.
Veröffentlicht: (2024)
von: Yarotsky, Dmitry, et al.
Veröffentlicht: (2024)
Does SGD really happen in tiny subspaces?
von: Song, Minhak, et al.
Veröffentlicht: (2024)
von: Song, Minhak, et al.
Veröffentlicht: (2024)
Sign-SGD via Parameter-Free Optimization
von: Medyakov, Daniil, et al.
Veröffentlicht: (2025)
von: Medyakov, Daniil, et al.
Veröffentlicht: (2025)
Can SGD Handle Heavy-Tailed Noise?
von: Fatkhullin, Ilyas, et al.
Veröffentlicht: (2025)
von: Fatkhullin, Ilyas, et al.
Veröffentlicht: (2025)
Optimal Growth Schedules for Batch Size and Learning Rate in SGD that Reduce SFO Complexity
von: Umeda, Hikaru, et al.
Veröffentlicht: (2025)
von: Umeda, Hikaru, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
On the Trajectories of SGD Without Replacement
von: Beneventano, Pierfrancesco
Veröffentlicht: (2023) -
Does Weight Decay Enhance Training Stability?
von: Saether, Marius, et al.
Veröffentlicht: (2026) -
Too Sharp, Too Sure: When Calibration Follows Curvature
von: Morosini, Alessandro, et al.
Veröffentlicht: (2026) -
Edge of Stochastic Stability: Revisiting the Edge of Stability for SGD
von: Andreyev, Arseniy, et al.
Veröffentlicht: (2024) -
Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias
von: Das, Mohua, et al.
Veröffentlicht: (2026)