ASGO: Adaptive Structured Gradient Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | An, Kang, Liu, Yuxing, Pan, Rui, Ren, Yi, Ma, Shiqian, Goldfarb, Donald, Zhang, Tong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Demystifying Manifold Constraints in LLM Pre-training
by: An, Kang, et al.
Published: (2026)
by: An, Kang, et al.
Published: (2026)
Accelerated Convergence of Stochastic Heavy Ball Method under Anisotropic Gradient Noise
by: Pan, Rui, et al.
Published: (2023)
by: Pan, Rui, et al.
Published: (2023)
AdaGrad under Anisotropic Smoothness
by: Liu, Yuxing, et al.
Published: (2024)
by: Liu, Yuxing, et al.
Published: (2024)
Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence
by: Liu, Yuxing, et al.
Published: (2025)
by: Liu, Yuxing, et al.
Published: (2025)
Riemannian Dueling Optimization
by: Ren, Yuxuan, et al.
Published: (2026)
by: Ren, Yuxuan, et al.
Published: (2026)
Unbiased Gradient Low-Rank Projection
by: Pan, Rui, et al.
Published: (2025)
by: Pan, Rui, et al.
Published: (2025)
StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models
by: Yu, Dingzhi, et al.
Published: (2026)
by: Yu, Dingzhi, et al.
Published: (2026)
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less
by: Liu, Yuxing, et al.
Published: (2026)
by: Liu, Yuxing, et al.
Published: (2026)
Stochastic Gradient Descent with Adaptive Data
by: Che, Ethan, et al.
Published: (2024)
by: Che, Ethan, et al.
Published: (2024)
An Adaptive and Parameter-Free Nesterov's Accelerated Gradient Method for Convex Optimization
by: Suh, Jaewook J., et al.
Published: (2025)
by: Suh, Jaewook J., et al.
Published: (2025)
Bregman Douglas-Rachford Splitting Method
by: Ma, Shiqian, et al.
Published: (2025)
by: Ma, Shiqian, et al.
Published: (2025)
Adaptive Gradient Normalization and Independent Sampling for (Stochastic) Generalized-Smooth Optimization
by: Yang, Yufeng, et al.
Published: (2024)
by: Yang, Yufeng, et al.
Published: (2024)
AutoBalance: An Automatic Balancing Framework for Training Physics-Informed Neural Networks
by: An, Kang, et al.
Published: (2025)
by: An, Kang, et al.
Published: (2025)
On the Complexity of Finite-Sum Smooth Optimization under the Polyak-Łojasiewicz Condition
by: Bai, Yunyan, et al.
Published: (2024)
by: Bai, Yunyan, et al.
Published: (2024)
On the Convergence of Adaptive Gradient Methods for Nonconvex Optimization
by: Zhou, Dongruo, et al.
Published: (2018)
by: Zhou, Dongruo, et al.
Published: (2018)
Decentralized and Equitable Optimal Transport
by: Lau, Ivan, et al.
Published: (2024)
by: Lau, Ivan, et al.
Published: (2024)
A New Inexact Proximal Linear Algorithm with Adaptive Stopping Criteria for Robust Phase Retrieval
by: Zheng, Zhong, et al.
Published: (2023)
by: Zheng, Zhong, et al.
Published: (2023)
Adaptive Moment Estimation Optimization Algorithm Using Projection Gradient for Deep Learning
by: Li, Yongqi, et al.
Published: (2025)
by: Li, Yongqi, et al.
Published: (2025)
Adaptive Optimization via Momentum on Variance-Normalized Gradients
by: Patitucci, Francisco, et al.
Published: (2026)
by: Patitucci, Francisco, et al.
Published: (2026)
ScaleBiO: Scalable Bilevel Optimization for LLM Data Reweighting
by: Pan, Rui, et al.
Published: (2024)
by: Pan, Rui, et al.
Published: (2024)
A Simple Adaptive Proximal Gradient Method for Nonconvex Optimization
by: Ye, Zilong, et al.
Published: (2025)
by: Ye, Zilong, et al.
Published: (2025)
Gradient-Variation Online Adaptivity for Accelerated Optimization with Hölder Smoothness
by: Zhao, Yuheng, et al.
Published: (2025)
by: Zhao, Yuheng, et al.
Published: (2025)
LDAdam: Adaptive Optimization from Low-Dimensional Gradient Statistics
by: Robert, Thomas, et al.
Published: (2024)
by: Robert, Thomas, et al.
Published: (2024)
Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness
by: Crawshaw, Michael, et al.
Published: (2025)
by: Crawshaw, Michael, et al.
Published: (2025)
On the Complexity of Decentralized Smooth Nonconvex Finite-Sum Optimization
by: Luo, Luo, et al.
Published: (2022)
by: Luo, Luo, et al.
Published: (2022)
Nesterov Finds GRAAL: Optimal and Adaptive Gradient Method for Convex Optimization
by: Borodich, Ekaterina, et al.
Published: (2025)
by: Borodich, Ekaterina, et al.
Published: (2025)
A Control Theoretic Framework for Adaptive Gradient Optimizers in Machine Learning
by: Chakrabarti, Kushal, et al.
Published: (2022)
by: Chakrabarti, Kushal, et al.
Published: (2022)
Interpreting Adaptive Gradient Methods by Parameter Scaling for Learning-Rate-Free Optimization
by: Suh, Min-Kook, et al.
Published: (2024)
by: Suh, Min-Kook, et al.
Published: (2024)
An Energy-Based Self-Adaptive Learning Rate for Stochastic Gradient Descent: Enhancing Unconstrained Optimization with VAV method
by: Zhang, Jiahao, et al.
Published: (2024)
by: Zhang, Jiahao, et al.
Published: (2024)
AdAdaGrad: Adaptive Batch Size Schemes for Adaptive Gradient Methods
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
Effective Bilevel Optimization via Minimax Reformulation
by: Wang, Xiaoyu, et al.
Published: (2023)
by: Wang, Xiaoyu, et al.
Published: (2023)
Adaptive Conditional Gradient Descent
by: Khademi, Abbas, et al.
Published: (2025)
by: Khademi, Abbas, et al.
Published: (2025)
Demystifying SGD with Doubly Stochastic Gradients
by: Kim, Kyurae, et al.
Published: (2024)
by: Kim, Kyurae, et al.
Published: (2024)
Multi-Objective Optimization via Wasserstein-Fisher-Rao Gradient Flow
by: Ren, Yinuo, et al.
Published: (2023)
by: Ren, Yinuo, et al.
Published: (2023)
Enhanced Adaptive Gradient Algorithms for Nonconvex-PL Minimax Optimization
by: Huang, Feihu, et al.
Published: (2023)
by: Huang, Feihu, et al.
Published: (2023)
Adaptive Proximal Gradient Method for Convex Optimization
by: Malitsky, Yura, et al.
Published: (2023)
by: Malitsky, Yura, et al.
Published: (2023)
A Single-Loop Algorithm for Decentralized Bilevel Optimization
by: Dong, Youran, et al.
Published: (2023)
by: Dong, Youran, et al.
Published: (2023)
Riemannian Zeroth-Order Gradient Estimation with Structure-Preserving Metrics for Geodesically Incomplete Manifolds
by: Ma, Shaocong, et al.
Published: (2026)
by: Ma, Shaocong, et al.
Published: (2026)
Reheated Gradient-based Discrete Sampling for Combinatorial Optimization
by: Li, Muheng, et al.
Published: (2025)
by: Li, Muheng, et al.
Published: (2025)
Towards Simple and Provable Parameter-Free Adaptive Gradient Methods
by: Tao, Yuanzhe, et al.
Published: (2024)
by: Tao, Yuanzhe, et al.
Published: (2024)
Similar Items
-
Demystifying Manifold Constraints in LLM Pre-training
by: An, Kang, et al.
Published: (2026) -
Accelerated Convergence of Stochastic Heavy Ball Method under Anisotropic Gradient Noise
by: Pan, Rui, et al.
Published: (2023) -
AdaGrad under Anisotropic Smoothness
by: Liu, Yuxing, et al.
Published: (2024) -
Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence
by: Liu, Yuxing, et al.
Published: (2025) -
Riemannian Dueling Optimization
by: Ren, Yuxuan, et al.
Published: (2026)