Zeroth-Order Optimization Finds Flat Minima
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Liang, Li, Bingcong, Thekumparampil, Kiran Koshy, Oh, Sewoong, Muehlebach, Michael, He, Niao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Zeroth-Order Optimization at the Edge of Stability
by: Song, Minhak, et al.
Published: (2026)
by: Song, Minhak, et al.
Published: (2026)
DPZero: Private Fine-Tuning of Language Models without Backpropagation
by: Zhang, Liang, et al.
Published: (2023)
by: Zhang, Liang, et al.
Published: (2023)
ConMeZO: Adaptive Descent-Direction Sampling for Gradient-Free Finetuning of Large Language Models
by: Behric, Lejs Deen, et al.
Published: (2025)
by: Behric, Lejs Deen, et al.
Published: (2025)
PoLAR: Polar-Decomposed Low-Rank Adapter Representation
by: Lion, Kai, et al.
Published: (2025)
by: Lion, Kai, et al.
Published: (2025)
Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe
by: Thekumparampil, Kiran Koshy, et al.
Published: (2024)
by: Thekumparampil, Kiran Koshy, et al.
Published: (2024)
Primal Methods for Variational Inequality Problems with Functional Constraints
by: Zhang, Liang, et al.
Published: (2024)
by: Zhang, Liang, et al.
Published: (2024)
On the Optimal Construction of Unbiased Gradient Estimators for Zeroth-Order Optimization
by: Ma, Shaocong, et al.
Published: (2025)
by: Ma, Shaocong, et al.
Published: (2025)
On the Benefits of Weight Normalization for Overparameterized Matrix Sensing
by: Wei, Yudong, et al.
Published: (2025)
by: Wei, Yudong, et al.
Published: (2025)
On the Crucial Role of Initialization for Matrix Factorization
by: Li, Bingcong, et al.
Published: (2024)
by: Li, Bingcong, et al.
Published: (2024)
Revisiting Zeroth-Order Optimization: Minimum-Variance Two-Point Estimators and Directionally Aligned Perturbations
by: Ma, Shaocong, et al.
Published: (2025)
by: Ma, Shaocong, et al.
Published: (2025)
A Sinkhorn-type Algorithm for Constrained Optimal Transport
by: Tang, Xun, et al.
Published: (2024)
by: Tang, Xun, et al.
Published: (2024)
Asynchronous Distributed Reinforcement Learning for LQR Control via Zeroth-Order Block Coordinate Descent
by: Jing, Gangshan, et al.
Published: (2021)
by: Jing, Gangshan, et al.
Published: (2021)
Accelerating Sinkhorn Algorithm with Sparse Newton Iterations
by: Tang, Xun, et al.
Published: (2024)
by: Tang, Xun, et al.
Published: (2024)
Characterizing the Training Dynamics of Private Fine-tuning with Langevin diffusion
by: Ke, Shuqi, et al.
Published: (2024)
by: Ke, Shuqi, et al.
Published: (2024)
On Constraints in First-Order Optimization: A View from Non-Smooth Dynamical Systems
by: Muehlebach, Michael, et al.
Published: (2021)
by: Muehlebach, Michael, et al.
Published: (2021)
Noise Stability Optimization for Finding Flat Minima: A Hessian-based Regularization Approach
by: Zhang, Hongyang R., et al.
Published: (2023)
by: Zhang, Hongyang R., et al.
Published: (2023)
Variance-reduced Zeroth-Order Methods for Fine-Tuning Language Models
by: Gautam, Tanmay, et al.
Published: (2024)
by: Gautam, Tanmay, et al.
Published: (2024)
Minimisation of Quasar-Convex Functions Using Random Zeroth-Order Oracles
by: Farzin, Amir Ali, et al.
Published: (2025)
by: Farzin, Amir Ali, et al.
Published: (2025)
Accelerated First-Order Optimization under Nonlinear Constraints
by: Muehlebach, Michael, et al.
Published: (2023)
by: Muehlebach, Michael, et al.
Published: (2023)
On Adaptivity in Zeroth-Order Optimization
by: Dbouk, Hassan, et al.
Published: (2026)
by: Dbouk, Hassan, et al.
Published: (2026)
Biased Stochastic First-Order Methods for Conditional Stochastic Optimization and Applications in Meta Learning
by: Hu, Yifan, et al.
Published: (2020)
by: Hu, Yifan, et al.
Published: (2020)
On Finding Small Hyper-Gradients in Bilevel Optimization: Hardness Results and Improved Analysis
by: Chen, Lesi, et al.
Published: (2023)
by: Chen, Lesi, et al.
Published: (2023)
Min-Max Optimisation for Nonconvex-Nonconcave Functions Using a Random Zeroth-Order Extragradient Algorithm
by: Farzin, Amir Ali, et al.
Published: (2025)
by: Farzin, Amir Ali, et al.
Published: (2025)
Optimal Guarantees for Algorithmic Reproducibility and Gradient Complexity in Convex Optimization
by: Zhang, Liang, et al.
Published: (2023)
by: Zhang, Liang, et al.
Published: (2023)
Primitive Agentic First-Order Optimization
by: Sala, R.
Published: (2024)
by: Sala, R.
Published: (2024)
TiAda: A Time-scale Adaptive Algorithm for Nonconvex Minimax Optimization
by: Li, Xiang, et al.
Published: (2022)
by: Li, Xiang, et al.
Published: (2022)
Decision-Dependent Stochastic Optimization: The Role of Distribution Dynamics
by: He, Zhiyu, et al.
Published: (2025)
by: He, Zhiyu, et al.
Published: (2025)
Private Zeroth-Order Nonsmooth Nonconvex Optimization
by: Zhang, Qinzi, et al.
Published: (2024)
by: Zhang, Qinzi, et al.
Published: (2024)
Zeroth-Order Methods for Stochastic Nonconvex Nonsmooth Composite Optimization
by: Chen, Ziyi, et al.
Published: (2025)
by: Chen, Ziyi, et al.
Published: (2025)
A Hessian-Aware Stochastic Differential Equation for Modelling SGD
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
Select-then-differentiate: Solving Bilevel Optimization with Manifold Lower-level Solution Sets
by: Masiha, Saeed, et al.
Published: (2026)
by: Masiha, Saeed, et al.
Published: (2026)
The Sample Complexity of Online Reinforcement Learning: A Multi-model Perspective
by: Muehlebach, Michael, et al.
Published: (2025)
by: Muehlebach, Michael, et al.
Published: (2025)
Certified Multi-Fidelity Zeroth-Order Optimization
by: de Montbrun, Étienne, et al.
Published: (2023)
by: de Montbrun, Étienne, et al.
Published: (2023)
Uncovering Symmetry Transfer in Large Language Models via Layer-Peeled Optimization
by: Du, Zhehang, et al.
Published: (2026)
by: Du, Zhehang, et al.
Published: (2026)
Gradient Descent with Polyak's Momentum Finds Flatter Minima via Large Catapults
by: Phunyaphibarn, Prin, et al.
Published: (2023)
by: Phunyaphibarn, Prin, et al.
Published: (2023)
Zeroth-Order Stochastic Mirror Descent Algorithms for Minimax Excess Risk Optimization
by: Gu, Zhihao, et al.
Published: (2024)
by: Gu, Zhihao, et al.
Published: (2024)
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less
by: Liu, Yuxing, et al.
Published: (2026)
by: Liu, Yuxing, et al.
Published: (2026)
On the Condition Number Dependency in Bilevel Optimization
by: Chen, Lesi, et al.
Published: (2025)
by: Chen, Lesi, et al.
Published: (2025)
Constructing Industrial-Scale Optimization Modeling Benchmark
by: Li, Zhong, et al.
Published: (2026)
by: Li, Zhong, et al.
Published: (2026)
Superquantile-Gibbs Relaxation for Minima-selection in Bilevel Optimization
by: Masiha, Saeed, et al.
Published: (2025)
by: Masiha, Saeed, et al.
Published: (2025)
Similar Items
-
Zeroth-Order Optimization at the Edge of Stability
by: Song, Minhak, et al.
Published: (2026) -
DPZero: Private Fine-Tuning of Language Models without Backpropagation
by: Zhang, Liang, et al.
Published: (2023) -
ConMeZO: Adaptive Descent-Direction Sampling for Gradient-Free Finetuning of Large Language Models
by: Behric, Lejs Deen, et al.
Published: (2025) -
PoLAR: Polar-Decomposed Low-Rank Adapter Representation
by: Lion, Kai, et al.
Published: (2025) -
Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe
by: Thekumparampil, Kiran Koshy, et al.
Published: (2024)