Training Diagonal Linear Networks with Stochastic Sharpness-Aware Minimization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Clara, Gabriel, Langer, Sophie, Schmidt-Hieber, Johannes |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Dropout Regularization Versus $\ell_2$-Penalization in the Linear Model
von: Clara, Gabriel, et al.
Veröffentlicht: (2023)
von: Clara, Gabriel, et al.
Veröffentlicht: (2023)
On the VC dimension of deep group convolutional neural networks
von: Sepliarskaia, Anna, et al.
Veröffentlicht: (2024)
von: Sepliarskaia, Anna, et al.
Veröffentlicht: (2024)
Spike-timing-dependent Hebbian learning as noisy gradient descent
von: Dexheimer, Niklas, et al.
Veröffentlicht: (2025)
von: Dexheimer, Niklas, et al.
Veröffentlicht: (2025)
Hyper Input Convex Neural Networks for Shape Constrained Learning and Optimal Transport
von: Hundrieser, Shayan, et al.
Veröffentlicht: (2026)
von: Hundrieser, Shayan, et al.
Veröffentlicht: (2026)
Asymptotics of Stochastic Gradient Descent with Dropout Regularization in Linear Models
von: Li, Jiaqi, et al.
Veröffentlicht: (2024)
von: Li, Jiaqi, et al.
Veröffentlicht: (2024)
High-dimensional Limit of SGD for Diagonal Linear Networks
von: Malaxechebarría, Begoña García, et al.
Veröffentlicht: (2026)
von: Malaxechebarría, Begoña García, et al.
Veröffentlicht: (2026)
Improving the Convergence Rates of Forward Gradient Descent with Repeated Sampling
von: Dexheimer, Niklas, et al.
Veröffentlicht: (2024)
von: Dexheimer, Niklas, et al.
Veröffentlicht: (2024)
Simplicity Bias via Global Convergence of Sharpness Minimization
von: Gatmiry, Khashayar, et al.
Veröffentlicht: (2024)
von: Gatmiry, Khashayar, et al.
Veröffentlicht: (2024)
A novel statistical approach to analyze image classification
von: Chen, Juntong, et al.
Veröffentlicht: (2022)
von: Chen, Juntong, et al.
Veröffentlicht: (2022)
Sharp Risk Bounds for Early-Stopping in Gaussian Linear Regression
von: Wegel, Tobias, et al.
Veröffentlicht: (2025)
von: Wegel, Tobias, et al.
Veröffentlicht: (2025)
Understanding the Effect of GCN Convolutions in Regression Tasks
von: Chen, Juntong, et al.
Veröffentlicht: (2024)
von: Chen, Juntong, et al.
Veröffentlicht: (2024)
Benign Overfitting without Linearity: Neural Network Classifiers Trained by Gradient Descent for Noisy Linear Data
von: Frei, Spencer, et al.
Veröffentlicht: (2022)
von: Frei, Spencer, et al.
Veröffentlicht: (2022)
Sharp Generalization for Nonparametric Regression in Interpolation Space by Over-Parameterized Neural Networks Trained with Preconditioned Gradient Descent and Early Stopping
von: Yang, Yingzhen, et al.
Veröffentlicht: (2024)
von: Yang, Yingzhen, et al.
Veröffentlicht: (2024)
Local convergence rates of the nonparametric least squares estimator with applications to transfer learning
von: Schmidt-Hieber, Johannes, et al.
Veröffentlicht: (2022)
von: Schmidt-Hieber, Johannes, et al.
Veröffentlicht: (2022)
Truncated LinUCB for Stochastic Linear Bandits
von: Song, Yanglei, et al.
Veröffentlicht: (2022)
von: Song, Yanglei, et al.
Veröffentlicht: (2022)
A Sharp KL-Convergence Analysis for Diffusion Models under Minimal Assumptions
von: Jain, Nishant, et al.
Veröffentlicht: (2025)
von: Jain, Nishant, et al.
Veröffentlicht: (2025)
Minimax Rate-Optimal Algorithms for High-Dimensional Stochastic Linear Bandits
von: Liu, Jingyu, et al.
Veröffentlicht: (2025)
von: Liu, Jingyu, et al.
Veröffentlicht: (2025)
Sharp Structure-Agnostic Lower Bounds for General Linear Functional Estimation
von: Jin, Jikai, et al.
Veröffentlicht: (2025)
von: Jin, Jikai, et al.
Veröffentlicht: (2025)
Optimal Excess Risk Bounds for Empirical Risk Minimization on $p$-Norm Linear Regression
von: Hanchi, Ayoub El, et al.
Veröffentlicht: (2023)
von: Hanchi, Ayoub El, et al.
Veröffentlicht: (2023)
Sharp Gaussian approximations for Decentralized Federated Learning
von: Bonnerjee, Soham, et al.
Veröffentlicht: (2025)
von: Bonnerjee, Soham, et al.
Veröffentlicht: (2025)
Understanding Learning Invariance in Deep Linear Networks
von: Duan, Hao, et al.
Veröffentlicht: (2025)
von: Duan, Hao, et al.
Veröffentlicht: (2025)
Stress-Aware Learning under KL Drift via Trust-Decayed Mirror Descent
von: Raj, Gabriel Nixon
Veröffentlicht: (2025)
von: Raj, Gabriel Nixon
Veröffentlicht: (2025)
Sharp Bounds for Poly-GNNs and the Effect of Graph Noise
von: Vinas, Luciano, et al.
Veröffentlicht: (2024)
von: Vinas, Luciano, et al.
Veröffentlicht: (2024)
Ordinal Patterns Based Change Points Detection
von: Betken, Annika, et al.
Veröffentlicht: (2025)
von: Betken, Annika, et al.
Veröffentlicht: (2025)
Multivariate Stochastic Dominance via Optimal Transport and Applications to Models Benchmarking
von: Rioux, Gabriel, et al.
Veröffentlicht: (2024)
von: Rioux, Gabriel, et al.
Veröffentlicht: (2024)
Affine Invariance in Continuous-Domain Convolutional Neural Networks
von: Mohaddes, Ali, et al.
Veröffentlicht: (2023)
von: Mohaddes, Ali, et al.
Veröffentlicht: (2023)
Sharp concentration of uniform generalization errors in binary linear classification
von: Nakakita, Shogo
Veröffentlicht: (2025)
von: Nakakita, Shogo
Veröffentlicht: (2025)
Sharp bounds on aggregate expert error
von: Kontorovich, Aryeh, et al.
Veröffentlicht: (2024)
von: Kontorovich, Aryeh, et al.
Veröffentlicht: (2024)
Sharp Inequalities between Total Variation and Hellinger Distances for Gaussian Mixtures
von: Jung, Joonhyuk, et al.
Veröffentlicht: (2026)
von: Jung, Joonhyuk, et al.
Veröffentlicht: (2026)
Belted and Ensembled Neural Network for Linear and Nonlinear Sufficient Dimension Reduction
von: Tang, Yin, et al.
Veröffentlicht: (2024)
von: Tang, Yin, et al.
Veröffentlicht: (2024)
Statistical Guarantees for Approximate Stationary Points of Shallow Neural Networks
von: Taheri, Mahsa, et al.
Veröffentlicht: (2022)
von: Taheri, Mahsa, et al.
Veröffentlicht: (2022)
The Adaptivity Barrier in Batched Nonparametric Bandits: Sharp Characterization of the Price of Unknown Margin
von: Jiang, Rong, et al.
Veröffentlicht: (2025)
von: Jiang, Rong, et al.
Veröffentlicht: (2025)
Sharp asymptotic theory for Q-learning with LDTZ learning rate and its generalization
von: Bonnerjee, Soham, et al.
Veröffentlicht: (2026)
von: Bonnerjee, Soham, et al.
Veröffentlicht: (2026)
Efficient Sampling with Discrete Diffusion Models: Sharp and Adaptive Guarantees
von: Dmitriev, Daniil, et al.
Veröffentlicht: (2026)
von: Dmitriev, Daniil, et al.
Veröffentlicht: (2026)
Minimax optimal submatrix detection: Sharp non-asymptotic rates
von: Knight, Parker, et al.
Veröffentlicht: (2026)
von: Knight, Parker, et al.
Veröffentlicht: (2026)
Unveil Conditional Diffusion Models with Classifier-free Guidance: A Sharp Statistical Theory
von: Fu, Hengyu, et al.
Veröffentlicht: (2024)
von: Fu, Hengyu, et al.
Veröffentlicht: (2024)
Zero-Order Sharpness-Aware Minimization
von: Fu, Yao, et al.
Veröffentlicht: (2025)
von: Fu, Yao, et al.
Veröffentlicht: (2025)
Chemical Reaction Networks Learn Better than Spiking Neural Networks
von: Jaffard, Sophie, et al.
Veröffentlicht: (2026)
von: Jaffard, Sophie, et al.
Veröffentlicht: (2026)
Revisiting Incremental Stochastic Majorization-Minimization Algorithms with Applications to Mixture of Experts
von: Tran, TrungKhang, et al.
Veröffentlicht: (2026)
von: Tran, TrungKhang, et al.
Veröffentlicht: (2026)
On the Variance, Admissibility, and Stability of Empirical Risk Minimization
von: Kur, Gil, et al.
Veröffentlicht: (2023)
von: Kur, Gil, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Dropout Regularization Versus $\ell_2$-Penalization in the Linear Model
von: Clara, Gabriel, et al.
Veröffentlicht: (2023) -
On the VC dimension of deep group convolutional neural networks
von: Sepliarskaia, Anna, et al.
Veröffentlicht: (2024) -
Spike-timing-dependent Hebbian learning as noisy gradient descent
von: Dexheimer, Niklas, et al.
Veröffentlicht: (2025) -
Hyper Input Convex Neural Networks for Shape Constrained Learning and Optimal Transport
von: Hundrieser, Shayan, et al.
Veröffentlicht: (2026) -
Asymptotics of Stochastic Gradient Descent with Dropout Regularization in Linear Models
von: Li, Jiaqi, et al.
Veröffentlicht: (2024)