Training Diagonal Linear Networks with Stochastic Sharpness-Aware Minimization
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Clara, Gabriel, Langer, Sophie, Schmidt-Hieber, Johannes |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Dropout Regularization Versus $\ell_2$-Penalization in the Linear Model
par: Clara, Gabriel, et autres
Publié: (2023)
par: Clara, Gabriel, et autres
Publié: (2023)
On the VC dimension of deep group convolutional neural networks
par: Sepliarskaia, Anna, et autres
Publié: (2024)
par: Sepliarskaia, Anna, et autres
Publié: (2024)
Spike-timing-dependent Hebbian learning as noisy gradient descent
par: Dexheimer, Niklas, et autres
Publié: (2025)
par: Dexheimer, Niklas, et autres
Publié: (2025)
Hyper Input Convex Neural Networks for Shape Constrained Learning and Optimal Transport
par: Hundrieser, Shayan, et autres
Publié: (2026)
par: Hundrieser, Shayan, et autres
Publié: (2026)
Asymptotics of Stochastic Gradient Descent with Dropout Regularization in Linear Models
par: Li, Jiaqi, et autres
Publié: (2024)
par: Li, Jiaqi, et autres
Publié: (2024)
High-dimensional Limit of SGD for Diagonal Linear Networks
par: Malaxechebarría, Begoña García, et autres
Publié: (2026)
par: Malaxechebarría, Begoña García, et autres
Publié: (2026)
Improving the Convergence Rates of Forward Gradient Descent with Repeated Sampling
par: Dexheimer, Niklas, et autres
Publié: (2024)
par: Dexheimer, Niklas, et autres
Publié: (2024)
Simplicity Bias via Global Convergence of Sharpness Minimization
par: Gatmiry, Khashayar, et autres
Publié: (2024)
par: Gatmiry, Khashayar, et autres
Publié: (2024)
A novel statistical approach to analyze image classification
par: Chen, Juntong, et autres
Publié: (2022)
par: Chen, Juntong, et autres
Publié: (2022)
Sharp Risk Bounds for Early-Stopping in Gaussian Linear Regression
par: Wegel, Tobias, et autres
Publié: (2025)
par: Wegel, Tobias, et autres
Publié: (2025)
Understanding the Effect of GCN Convolutions in Regression Tasks
par: Chen, Juntong, et autres
Publié: (2024)
par: Chen, Juntong, et autres
Publié: (2024)
Benign Overfitting without Linearity: Neural Network Classifiers Trained by Gradient Descent for Noisy Linear Data
par: Frei, Spencer, et autres
Publié: (2022)
par: Frei, Spencer, et autres
Publié: (2022)
Sharp Generalization for Nonparametric Regression in Interpolation Space by Over-Parameterized Neural Networks Trained with Preconditioned Gradient Descent and Early Stopping
par: Yang, Yingzhen, et autres
Publié: (2024)
par: Yang, Yingzhen, et autres
Publié: (2024)
Local convergence rates of the nonparametric least squares estimator with applications to transfer learning
par: Schmidt-Hieber, Johannes, et autres
Publié: (2022)
par: Schmidt-Hieber, Johannes, et autres
Publié: (2022)
Truncated LinUCB for Stochastic Linear Bandits
par: Song, Yanglei, et autres
Publié: (2022)
par: Song, Yanglei, et autres
Publié: (2022)
A Sharp KL-Convergence Analysis for Diffusion Models under Minimal Assumptions
par: Jain, Nishant, et autres
Publié: (2025)
par: Jain, Nishant, et autres
Publié: (2025)
Minimax Rate-Optimal Algorithms for High-Dimensional Stochastic Linear Bandits
par: Liu, Jingyu, et autres
Publié: (2025)
par: Liu, Jingyu, et autres
Publié: (2025)
Sharp Structure-Agnostic Lower Bounds for General Linear Functional Estimation
par: Jin, Jikai, et autres
Publié: (2025)
par: Jin, Jikai, et autres
Publié: (2025)
Optimal Excess Risk Bounds for Empirical Risk Minimization on $p$-Norm Linear Regression
par: Hanchi, Ayoub El, et autres
Publié: (2023)
par: Hanchi, Ayoub El, et autres
Publié: (2023)
Sharp Gaussian approximations for Decentralized Federated Learning
par: Bonnerjee, Soham, et autres
Publié: (2025)
par: Bonnerjee, Soham, et autres
Publié: (2025)
Understanding Learning Invariance in Deep Linear Networks
par: Duan, Hao, et autres
Publié: (2025)
par: Duan, Hao, et autres
Publié: (2025)
Stress-Aware Learning under KL Drift via Trust-Decayed Mirror Descent
par: Raj, Gabriel Nixon
Publié: (2025)
par: Raj, Gabriel Nixon
Publié: (2025)
Sharp Bounds for Poly-GNNs and the Effect of Graph Noise
par: Vinas, Luciano, et autres
Publié: (2024)
par: Vinas, Luciano, et autres
Publié: (2024)
Ordinal Patterns Based Change Points Detection
par: Betken, Annika, et autres
Publié: (2025)
par: Betken, Annika, et autres
Publié: (2025)
Multivariate Stochastic Dominance via Optimal Transport and Applications to Models Benchmarking
par: Rioux, Gabriel, et autres
Publié: (2024)
par: Rioux, Gabriel, et autres
Publié: (2024)
Affine Invariance in Continuous-Domain Convolutional Neural Networks
par: Mohaddes, Ali, et autres
Publié: (2023)
par: Mohaddes, Ali, et autres
Publié: (2023)
Sharp concentration of uniform generalization errors in binary linear classification
par: Nakakita, Shogo
Publié: (2025)
par: Nakakita, Shogo
Publié: (2025)
Sharp bounds on aggregate expert error
par: Kontorovich, Aryeh, et autres
Publié: (2024)
par: Kontorovich, Aryeh, et autres
Publié: (2024)
Sharp Inequalities between Total Variation and Hellinger Distances for Gaussian Mixtures
par: Jung, Joonhyuk, et autres
Publié: (2026)
par: Jung, Joonhyuk, et autres
Publié: (2026)
Belted and Ensembled Neural Network for Linear and Nonlinear Sufficient Dimension Reduction
par: Tang, Yin, et autres
Publié: (2024)
par: Tang, Yin, et autres
Publié: (2024)
Statistical Guarantees for Approximate Stationary Points of Shallow Neural Networks
par: Taheri, Mahsa, et autres
Publié: (2022)
par: Taheri, Mahsa, et autres
Publié: (2022)
The Adaptivity Barrier in Batched Nonparametric Bandits: Sharp Characterization of the Price of Unknown Margin
par: Jiang, Rong, et autres
Publié: (2025)
par: Jiang, Rong, et autres
Publié: (2025)
Sharp asymptotic theory for Q-learning with LDTZ learning rate and its generalization
par: Bonnerjee, Soham, et autres
Publié: (2026)
par: Bonnerjee, Soham, et autres
Publié: (2026)
Efficient Sampling with Discrete Diffusion Models: Sharp and Adaptive Guarantees
par: Dmitriev, Daniil, et autres
Publié: (2026)
par: Dmitriev, Daniil, et autres
Publié: (2026)
Minimax optimal submatrix detection: Sharp non-asymptotic rates
par: Knight, Parker, et autres
Publié: (2026)
par: Knight, Parker, et autres
Publié: (2026)
Unveil Conditional Diffusion Models with Classifier-free Guidance: A Sharp Statistical Theory
par: Fu, Hengyu, et autres
Publié: (2024)
par: Fu, Hengyu, et autres
Publié: (2024)
Zero-Order Sharpness-Aware Minimization
par: Fu, Yao, et autres
Publié: (2025)
par: Fu, Yao, et autres
Publié: (2025)
Chemical Reaction Networks Learn Better than Spiking Neural Networks
par: Jaffard, Sophie, et autres
Publié: (2026)
par: Jaffard, Sophie, et autres
Publié: (2026)
Revisiting Incremental Stochastic Majorization-Minimization Algorithms with Applications to Mixture of Experts
par: Tran, TrungKhang, et autres
Publié: (2026)
par: Tran, TrungKhang, et autres
Publié: (2026)
On the Variance, Admissibility, and Stability of Empirical Risk Minimization
par: Kur, Gil, et autres
Publié: (2023)
par: Kur, Gil, et autres
Publié: (2023)
Documents similaires
-
Dropout Regularization Versus $\ell_2$-Penalization in the Linear Model
par: Clara, Gabriel, et autres
Publié: (2023) -
On the VC dimension of deep group convolutional neural networks
par: Sepliarskaia, Anna, et autres
Publié: (2024) -
Spike-timing-dependent Hebbian learning as noisy gradient descent
par: Dexheimer, Niklas, et autres
Publié: (2025) -
Hyper Input Convex Neural Networks for Shape Constrained Learning and Optimal Transport
par: Hundrieser, Shayan, et autres
Publié: (2026) -
Asymptotics of Stochastic Gradient Descent with Dropout Regularization in Linear Models
par: Li, Jiaqi, et autres
Publié: (2024)