Simplicity Bias via Global Convergence of Sharpness Minimization
Fuente:
arXiv
Saved in:
| Main Authors: | Gatmiry, Khashayar, Li, Zhiyuan, Reddi, Sashank J., Jegelka, Stefanie |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the Role of Depth and Looping for In-Context Learning with Task Diversity
by: Gatmiry, Khashayar, et al.
Published: (2024)
by: Gatmiry, Khashayar, et al.
Published: (2024)
Computing Optimal Regularizers for Online Linear Optimization
by: Gatmiry, Khashayar, et al.
Published: (2024)
by: Gatmiry, Khashayar, et al.
Published: (2024)
Can Looped Transformers Learn to Implement Multi-step Gradient Descent for In-context Learning?
by: Gatmiry, Khashayar, et al.
Published: (2024)
by: Gatmiry, Khashayar, et al.
Published: (2024)
High-accuracy and dimension-free sampling with diffusions
by: Gatmiry, Khashayar, et al.
Published: (2026)
by: Gatmiry, Khashayar, et al.
Published: (2026)
Learning Mixtures of Gaussians Using Diffusion Models
by: Gatmiry, Khashayar, et al.
Published: (2024)
by: Gatmiry, Khashayar, et al.
Published: (2024)
Near-Optimal Algorithms for Group Distributionally Robust Optimization and Beyond
by: Soma, Tasuku, et al.
Published: (2022)
by: Soma, Tasuku, et al.
Published: (2022)
A Sharp KL-Convergence Analysis for Diffusion Models under Minimal Assumptions
by: Jain, Nishant, et al.
Published: (2025)
by: Jain, Nishant, et al.
Published: (2025)
Principled Out-of-Distribution Generalization via Simplicity
by: Ge, Jiawei, et al.
Published: (2025)
by: Ge, Jiawei, et al.
Published: (2025)
Training Diagonal Linear Networks with Stochastic Sharpness-Aware Minimization
by: Clara, Gabriel, et al.
Published: (2025)
by: Clara, Gabriel, et al.
Published: (2025)
On the hardness of learning under symmetries
by: Kiani, Bobak T., et al.
Published: (2024)
by: Kiani, Bobak T., et al.
Published: (2024)
Adversarial Online Learning with Temporal Feedback Graphs
by: Gatmiry, Khashayar, et al.
Published: (2024)
by: Gatmiry, Khashayar, et al.
Published: (2024)
O(d/T) Convergence Theory for Diffusion Probabilistic Models under Minimal Assumptions
by: Li, Gen, et al.
Published: (2024)
by: Li, Gen, et al.
Published: (2024)
A Sharp Convergence Theory for The Probability Flow ODEs of Diffusion Models
by: Li, Gen, et al.
Published: (2024)
by: Li, Gen, et al.
Published: (2024)
Convergence of Statistical Estimators via Mutual Information Bounds
by: Khribch, El Mahdi, et al.
Published: (2024)
by: Khribch, El Mahdi, et al.
Published: (2024)
Sharp Gaussian approximations for Decentralized Federated Learning
by: Bonnerjee, Soham, et al.
Published: (2025)
by: Bonnerjee, Soham, et al.
Published: (2025)
Sharp Bounds for Poly-GNNs and the Effect of Graph Noise
by: Vinas, Luciano, et al.
Published: (2024)
by: Vinas, Luciano, et al.
Published: (2024)
Global Convergence in Training Large-Scale Transformers
by: Gao, Cheng, et al.
Published: (2024)
by: Gao, Cheng, et al.
Published: (2024)
Improved Convergence of Score-Based Diffusion Models via Prediction-Correction
by: Pedrotti, Francesco, et al.
Published: (2023)
by: Pedrotti, Francesco, et al.
Published: (2023)
Optimal Convergence Analysis of DDPM for General Distributions
by: Jiao, Yuchen, et al.
Published: (2025)
by: Jiao, Yuchen, et al.
Published: (2025)
Sharp concentration of uniform generalization errors in binary linear classification
by: Nakakita, Shogo
Published: (2025)
by: Nakakita, Shogo
Published: (2025)
Sharp Risk Bounds for Early-Stopping in Gaussian Linear Regression
by: Wegel, Tobias, et al.
Published: (2025)
by: Wegel, Tobias, et al.
Published: (2025)
Sharp detection of low-dimensional structure in probability measures via dimensional logarithmic Sobolev inequalities
by: Li, Matthew T. C., et al.
Published: (2024)
by: Li, Matthew T. C., et al.
Published: (2024)
Sharp bounds on aggregate expert error
by: Kontorovich, Aryeh, et al.
Published: (2024)
by: Kontorovich, Aryeh, et al.
Published: (2024)
Sharp Inequalities between Total Variation and Hellinger Distances for Gaussian Mixtures
by: Jung, Joonhyuk, et al.
Published: (2026)
by: Jung, Joonhyuk, et al.
Published: (2026)
Sharp Generalization for Nonparametric Regression in Interpolation Space by Over-Parameterized Neural Networks Trained with Preconditioned Gradient Descent and Early Stopping
by: Yang, Yingzhen, et al.
Published: (2024)
by: Yang, Yingzhen, et al.
Published: (2024)
Residual-as-Teacher: Mitigating Bias Propagation in Student--Teacher Estimation
by: Yamamoto, Kakei, et al.
Published: (2026)
by: Yamamoto, Kakei, et al.
Published: (2026)
Sharp asymptotic theory for Q-learning with LDTZ learning rate and its generalization
by: Bonnerjee, Soham, et al.
Published: (2026)
by: Bonnerjee, Soham, et al.
Published: (2026)
The Adaptivity Barrier in Batched Nonparametric Bandits: Sharp Characterization of the Price of Unknown Margin
by: Jiang, Rong, et al.
Published: (2025)
by: Jiang, Rong, et al.
Published: (2025)
Efficient Sampling with Discrete Diffusion Models: Sharp and Adaptive Guarantees
by: Dmitriev, Daniil, et al.
Published: (2026)
by: Dmitriev, Daniil, et al.
Published: (2026)
Minimax optimal submatrix detection: Sharp non-asymptotic rates
by: Knight, Parker, et al.
Published: (2026)
by: Knight, Parker, et al.
Published: (2026)
Dimension-Free Convergence of Discrete Diffusion Models: Adjoint Equations Induce the Right Space
by: Kan, Kelvin, et al.
Published: (2026)
by: Kan, Kelvin, et al.
Published: (2026)
Unveil Conditional Diffusion Models with Classifier-free Guidance: A Sharp Statistical Theory
by: Fu, Hengyu, et al.
Published: (2024)
by: Fu, Hengyu, et al.
Published: (2024)
Targeted Separation and Convergence with Kernel Discrepancies
by: Barp, Alessandro, et al.
Published: (2022)
by: Barp, Alessandro, et al.
Published: (2022)
Wasserstein Convergence of Critically Damped Langevin Diffusions
by: Strasman, Stanislas, et al.
Published: (2025)
by: Strasman, Stanislas, et al.
Published: (2025)
Towards Faster Non-Asymptotic Convergence for Diffusion-Based Generative Models
by: Li, Gen, et al.
Published: (2023)
by: Li, Gen, et al.
Published: (2023)
On the Variance, Admissibility, and Stability of Empirical Risk Minimization
by: Kur, Gil, et al.
Published: (2023)
by: Kur, Gil, et al.
Published: (2023)
A Researcher's Guide to Empirical Risk Minimization
by: van der Laan, Lars
Published: (2026)
by: van der Laan, Lars
Published: (2026)
Fast Rates for Nonstationary Weighted Risk Minimization
by: Brock, Tobias, et al.
Published: (2026)
by: Brock, Tobias, et al.
Published: (2026)
Linear Convergence of Diffusion Models Under the Manifold Hypothesis
by: Potaptchik, Peter, et al.
Published: (2024)
by: Potaptchik, Peter, et al.
Published: (2024)
Statistical Convergence of Spherical First Hitting Diffusion Models
by: Bienewald, Simon, et al.
Published: (2026)
by: Bienewald, Simon, et al.
Published: (2026)
Similar Items
-
On the Role of Depth and Looping for In-Context Learning with Task Diversity
by: Gatmiry, Khashayar, et al.
Published: (2024) -
Computing Optimal Regularizers for Online Linear Optimization
by: Gatmiry, Khashayar, et al.
Published: (2024) -
Can Looped Transformers Learn to Implement Multi-step Gradient Descent for In-context Learning?
by: Gatmiry, Khashayar, et al.
Published: (2024) -
High-accuracy and dimension-free sampling with diffusions
by: Gatmiry, Khashayar, et al.
Published: (2026) -
Learning Mixtures of Gaussians Using Diffusion Models
by: Gatmiry, Khashayar, et al.
Published: (2024)