How to escape sharp minima with random perturbations
Fuente:
arXiv
Saved in:
| Main Authors: | Ahn, Kwangjun, Jadbabaie, Ali, Sra, Suvrit |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Linear attention is (maybe) all you need (to understand transformer optimization)
by: Ahn, Kwangjun, et al.
Published: (2023)
by: Ahn, Kwangjun, et al.
Published: (2023)
Riemannian Bilevel Optimization
by: Dutta, Sanchayan, et al.
Published: (2024)
by: Dutta, Sanchayan, et al.
Published: (2024)
Through the River: Understanding the Benefit of Schedule-Free Methods for Language Model Training
by: Song, Minhak, et al.
Published: (2025)
by: Song, Minhak, et al.
Published: (2025)
Toward generalizable learning of all (linear) first-order methods via memory augmented Transformers
by: Dutta, Sanchayan, et al.
Published: (2024)
by: Dutta, Sanchayan, et al.
Published: (2024)
Linearly Convergent Algorithms for Nonsmooth Problems with Unknown Smooth Pieces
by: Zhang, Zhe, et al.
Published: (2025)
by: Zhang, Zhe, et al.
Published: (2025)
Implicit Bias in Matrix Factorization and its Explicit Realization in a New Architecture
by: Hou, Yikun, et al.
Published: (2025)
by: Hou, Yikun, et al.
Published: (2025)
Dion: Distributed Orthonormalized Updates
by: Ahn, Kwangjun, et al.
Published: (2025)
by: Ahn, Kwangjun, et al.
Published: (2025)
Adam with model exponential moving average is effective for nonconvex optimization
by: Ahn, Kwangjun, et al.
Published: (2024)
by: Ahn, Kwangjun, et al.
Published: (2024)
Cost-Driven Representation Learning for Linear Quadratic Gaussian Control: Part I
by: Tian, Yi, et al.
Published: (2022)
by: Tian, Yi, et al.
Published: (2022)
Cost-Driven Representation Learning for Linear Quadratic Gaussian Control: Part II
by: Tian, Yi, et al.
Published: (2026)
by: Tian, Yi, et al.
Published: (2026)
Tight Generalization Bounds for Noiseless Inverse Optimization
by: Fatemi, Pouria, et al.
Published: (2026)
by: Fatemi, Pouria, et al.
Published: (2026)
First-Order Methods for Linearly Constrained Bilevel Optimization
by: Kornowski, Guy, et al.
Published: (2024)
by: Kornowski, Guy, et al.
Published: (2024)
General framework for online-to-nonconvex conversion: Schedule-free SGD is also effective for nonconvex optimization
by: Ahn, Kwangjun, et al.
Published: (2024)
by: Ahn, Kwangjun, et al.
Published: (2024)
Does SGD really happen in tiny subspaces?
by: Song, Minhak, et al.
Published: (2024)
by: Song, Minhak, et al.
Published: (2024)
Revisiting Frank-Wolfe for Structured Nonconvex Optimization
by: Maskan, Hoomaan, et al.
Published: (2025)
by: Maskan, Hoomaan, et al.
Published: (2025)
Understanding Adam Optimizer via Online Learning of Updates: Adam is FTRL in Disguise
by: Ahn, Kwangjun, et al.
Published: (2024)
by: Ahn, Kwangjun, et al.
Published: (2024)
Online Learning for Supervisory Switching Control
by: Sun, Haoyuan, et al.
Published: (2026)
by: Sun, Haoyuan, et al.
Published: (2026)
A least-square method for non-asymptotic identification in linear switching control
by: Sun, Haoyuan, et al.
Published: (2024)
by: Sun, Haoyuan, et al.
Published: (2024)
How Memory in Optimization Algorithms Implicitly Modifies the Loss
by: Cattaneo, Matias D., et al.
Published: (2025)
by: Cattaneo, Matias D., et al.
Published: (2025)
DualSchool: How Reliable are LLMs for Optimization Education?
by: Klamkin, Michael, et al.
Published: (2025)
by: Klamkin, Michael, et al.
Published: (2025)
How Well Can Transformers Emulate In-context Newton's Method?
by: Giannou, Angeliki, et al.
Published: (2024)
by: Giannou, Angeliki, et al.
Published: (2024)
How Does Critical Batch Size Scale in Pre-training?
by: Zhang, Hanlin, et al.
Published: (2024)
by: Zhang, Hanlin, et al.
Published: (2024)
Personalized Multi-tier Federated Learning
by: Banerjee, Sourasekhar, et al.
Published: (2024)
by: Banerjee, Sourasekhar, et al.
Published: (2024)
Graph Transformers Dream of Electric Flow
by: Cheng, Xiang, et al.
Published: (2024)
by: Cheng, Xiang, et al.
Published: (2024)
A Median Perspective on Unlabeled Data for Out-of-Distribution Detection
by: Abbas, Momin, et al.
Published: (2025)
by: Abbas, Momin, et al.
Published: (2025)
Trees to Flows and Back: Unifying Decision Trees and Diffusion Models
by: Ramachandran, Sai Niranjan, et al.
Published: (2026)
by: Ramachandran, Sai Niranjan, et al.
Published: (2026)
Federated Optimization of Smooth Loss Functions
by: Jadbabaie, Ali, et al.
Published: (2022)
by: Jadbabaie, Ali, et al.
Published: (2022)
Seeing Through Risk: A Symbolic Approximation of Prospect Theory
by: Yousaf, Ali Arslan, et al.
Published: (2025)
by: Yousaf, Ali Arslan, et al.
Published: (2025)
Adaptive Smooth Tchebycheff Attention for Multi-Objective Policy Optimization
by: Murillo-Gonzalez, Alejandro, et al.
Published: (2026)
by: Murillo-Gonzalez, Alejandro, et al.
Published: (2026)
When and How Unlabeled Data Provably Improve In-Context Learning
by: Li, Yingcong, et al.
Published: (2025)
by: Li, Yingcong, et al.
Published: (2025)
On the Interplay Between Stepsize Tuning and Progressive Sharpening
by: Roulet, Vincent, et al.
Published: (2023)
by: Roulet, Vincent, et al.
Published: (2023)
TaskMet: Task-Driven Metric Learning for Model Learning
by: Bansal, Dishank, et al.
Published: (2023)
by: Bansal, Dishank, et al.
Published: (2023)
Multi-Objective Optimization for Sparse Deep Multi-Task Learning
by: Hotegni, S. S., et al.
Published: (2023)
by: Hotegni, S. S., et al.
Published: (2023)
Online Submodular Maximization via Online Convex Optimization
by: Salem, Tareq Si, et al.
Published: (2023)
by: Salem, Tareq Si, et al.
Published: (2023)
Comparative Analysis of Optimization Strategies for K-means Clustering in Big Data Contexts: A Review
by: Mussabayev, Ravil, et al.
Published: (2023)
by: Mussabayev, Ravil, et al.
Published: (2023)
A multiobjective continuation method to compute the regularization path of deep neural networks
by: Amakor, Augustina C., et al.
Published: (2023)
by: Amakor, Augustina C., et al.
Published: (2023)
On Finding Small Hyper-Gradients in Bilevel Optimization: Hardness Results and Improved Analysis
by: Chen, Lesi, et al.
Published: (2023)
by: Chen, Lesi, et al.
Published: (2023)
Federated Distributionally Robust Optimization with Non-Convex Objectives: Algorithm and Analysis
by: Jiao, Yang, et al.
Published: (2023)
by: Jiao, Yang, et al.
Published: (2023)
Nash Equilibria, Regularization and Computation in Optimal Transport-Based Distributionally Robust Optimization
by: Shafiee, Soroosh, et al.
Published: (2023)
by: Shafiee, Soroosh, et al.
Published: (2023)
An improved column-generation-based matheuristic for learning classification trees
by: Patel, Krunal Kishor, et al.
Published: (2023)
by: Patel, Krunal Kishor, et al.
Published: (2023)
Similar Items
-
Linear attention is (maybe) all you need (to understand transformer optimization)
by: Ahn, Kwangjun, et al.
Published: (2023) -
Riemannian Bilevel Optimization
by: Dutta, Sanchayan, et al.
Published: (2024) -
Through the River: Understanding the Benefit of Schedule-Free Methods for Language Model Training
by: Song, Minhak, et al.
Published: (2025) -
Toward generalizable learning of all (linear) first-order methods via memory augmented Transformers
by: Dutta, Sanchayan, et al.
Published: (2024) -
Linearly Convergent Algorithms for Nonsmooth Problems with Unknown Smooth Pieces
by: Zhang, Zhe, et al.
Published: (2025)