Optimization Hyper-parameter Laws for Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Xie, Xingyu, Ding, Kuangyu, Yan, Shuicheng, Toh, Kim-Chuan, Wei, Tianwen |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On exploration of an interior mirror descent flow for stochastic nonconvex constrained problem
by: Ding, Kuangyu, et al.
Published: (2025)
by: Ding, Kuangyu, et al.
Published: (2025)
Nonconvex Stochastic Bregman Proximal Gradient Method with Application to Deep Learning
by: Ding, Kuangyu, et al.
Published: (2023)
by: Ding, Kuangyu, et al.
Published: (2023)
Developing Lagrangian-based Methods for Nonsmooth Nonconvex Optimization
by: Xiao, Nachuan, et al.
Published: (2024)
by: Xiao, Nachuan, et al.
Published: (2024)
LoCo: Low-Bit Communication Adaptor for Large-scale Model Training
by: Xie, Xingyu, et al.
Published: (2024)
by: Xie, Xingyu, et al.
Published: (2024)
Stochastic Bregman Subgradient Methods for Nonsmooth Nonconvex Optimization Problems
by: Ding, Kuangyu, et al.
Published: (2024)
by: Ding, Kuangyu, et al.
Published: (2024)
Memory-Efficient 4-bit Preconditioned Stochastic Optimization
by: Li, Jingyang, et al.
Published: (2024)
by: Li, Jingyang, et al.
Published: (2024)
Adan: Adaptive Nesterov Momentum Algorithm for Faster Optimizing Deep Models
by: Xie, Xingyu, et al.
Published: (2022)
by: Xie, Xingyu, et al.
Published: (2022)
Adam-family Methods for Nonsmooth Optimization with Convergence Guarantees
by: Xiao, Nachuan, et al.
Published: (2023)
by: Xiao, Nachuan, et al.
Published: (2023)
Stochastic Subgradient Methods with Guaranteed Global Stability in Nonsmooth Nonconvex Optimization
by: Xiao, Nachuan, et al.
Published: (2023)
by: Xiao, Nachuan, et al.
Published: (2023)
Learning Graph Laplacian with MCP
by: Zhang, Yangjing, et al.
Published: (2020)
by: Zhang, Yangjing, et al.
Published: (2020)
Accelerating nuclear-norm regularized low-rank matrix optimization through Burer-Monteiro decomposition
by: Lee, Ching-pei, et al.
Published: (2022)
by: Lee, Ching-pei, et al.
Published: (2022)
A Multi-objective Newton Optimization Algorithm for Hyper-Parameter Search
by: Xu, Qinwu
Published: (2024)
by: Xu, Qinwu
Published: (2024)
Subspace Optimization for Large Language Models with Convergence Guarantees
by: He, Yutong, et al.
Published: (2024)
by: He, Yutong, et al.
Published: (2024)
Tractable hierarchies of convex relaxations for polynomial optimization on the nonnegative orthant
by: Mai, Ngoc Hoang Anh, et al.
Published: (2022)
by: Mai, Ngoc Hoang Anh, et al.
Published: (2022)
DOVA-PATBM: An Intelligent, Adaptive, and Scalable Framework for Optimizing Large-Scale EV Charging Infrastructure
by: Li, Chuan, et al.
Published: (2025)
by: Li, Chuan, et al.
Published: (2025)
Optimal Rates for Robust Stochastic Convex Optimization
by: Gao, Changyu, et al.
Published: (2024)
by: Gao, Changyu, et al.
Published: (2024)
Scaling Laws for Gradient Descent and Sign Descent for Linear Bigram Models under Zipf's Law
by: Kunstner, Frederik, et al.
Published: (2025)
by: Kunstner, Frederik, et al.
Published: (2025)
LiMuon: Light and Fast Muon Optimizer for Large Models
by: Huang, Feihu, et al.
Published: (2025)
by: Huang, Feihu, et al.
Published: (2025)
An Inexact Halpern Iteration with Application to Distributionally Robust Optimization
by: Liang, Ling, et al.
Published: (2024)
by: Liang, Ling, et al.
Published: (2024)
PDHG-Unrolled Learning-to-Optimize Method for Large-Scale Linear Programming
by: Li, Bingheng, et al.
Published: (2024)
by: Li, Bingheng, et al.
Published: (2024)
On Finding Small Hyper-Gradients in Bilevel Optimization: Hardness Results and Improved Analysis
by: Chen, Lesi, et al.
Published: (2023)
by: Chen, Lesi, et al.
Published: (2023)
Inexact Bregman Proximal Gradient Method and its Inertial Variant with Absolute and Partial Relative Stopping Criteria
by: Yang, Lei, et al.
Published: (2021)
by: Yang, Lei, et al.
Published: (2021)
Dynamic Regret via Discounted-to-Dynamic Reduction with Applications to Curved Losses and Adam Optimizer
by: Xie, Yan-Feng, et al.
Published: (2026)
by: Xie, Yan-Feng, et al.
Published: (2026)
A Hyper-Transformer model for Controllable Pareto Front Learning with Split Feasibility Constraints
by: Tuan, Tran Anh, et al.
Published: (2024)
by: Tuan, Tran Anh, et al.
Published: (2024)
Robust principal component analysis with rank and cardinality regularization under matrix factorization
by: Li, Wenjing, et al.
Published: (2026)
by: Li, Wenjing, et al.
Published: (2026)
Abide by the Law and Follow the Flow: Conservation Laws for Gradient Flows
by: Marcotte, Sibylle, et al.
Published: (2023)
by: Marcotte, Sibylle, et al.
Published: (2023)
From Large Language Models and Optimization to Decision Optimization CoPilot: A Research Manifesto
by: Wasserkrug, Segev, et al.
Published: (2024)
by: Wasserkrug, Segev, et al.
Published: (2024)
Sinkhorn Distributionally Robust Optimization
by: Wang, Jie, et al.
Published: (2021)
by: Wang, Jie, et al.
Published: (2021)
Stability and Generalization for Stochastic Recursive Momentum-based Algorithms for (Strongly-)Convex One to $K$-Level Stochastic Optimizations
by: Pan, Xiaokang, et al.
Published: (2024)
by: Pan, Xiaokang, et al.
Published: (2024)
NewVEM: A Newton Vertex Exchange Method for a Class of Constrained Self-Concordant Minimization Problems
by: Liang, Ling, et al.
Published: (2024)
by: Liang, Ling, et al.
Published: (2024)
Self-supervised Equality Embedded Deep Lagrange Dual for Approximate Constrained Optimization
by: Kim, Minsoo, et al.
Published: (2023)
by: Kim, Minsoo, et al.
Published: (2023)
Partial Envelope for Optimization Problem with Nonconvex Constraints
by: Hu, Xiaoyin, et al.
Published: (2025)
by: Hu, Xiaoyin, et al.
Published: (2025)
GNMR: Runtime Stability Control for Low-Precision Large Language Model Training
by: Kong, Boao, et al.
Published: (2026)
by: Kong, Boao, et al.
Published: (2026)
The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training
by: Schaipp, Fabian, et al.
Published: (2025)
by: Schaipp, Fabian, et al.
Published: (2025)
Uncovering Symmetry Transfer in Large Language Models via Layer-Peeled Optimization
by: Du, Zhehang, et al.
Published: (2026)
by: Du, Zhehang, et al.
Published: (2026)
Implicit Bias of AdamW: $\ell_\infty$ Norm Constrained Optimization
by: Xie, Shuo, et al.
Published: (2024)
by: Xie, Shuo, et al.
Published: (2024)
Distributionally Robust Optimization via Iterative Algorithms in Continuous Probability Spaces
by: Zhu, Linglingzhi, et al.
Published: (2024)
by: Zhu, Linglingzhi, et al.
Published: (2024)
OptScaler: A Collaborative Framework for Robust Autoscaling in the Cloud
by: Zou, Ding, et al.
Published: (2023)
by: Zou, Ding, et al.
Published: (2023)
PAC-Bayes Meets Online Contextual Optimization
by: Xie, Zhuojun, et al.
Published: (2025)
by: Xie, Zhuojun, et al.
Published: (2025)
A Minimization Approach for Minimax Optimization with Coupled Constraints
by: Hu, Xiaoyin, et al.
Published: (2024)
by: Hu, Xiaoyin, et al.
Published: (2024)
Similar Items
-
On exploration of an interior mirror descent flow for stochastic nonconvex constrained problem
by: Ding, Kuangyu, et al.
Published: (2025) -
Nonconvex Stochastic Bregman Proximal Gradient Method with Application to Deep Learning
by: Ding, Kuangyu, et al.
Published: (2023) -
Developing Lagrangian-based Methods for Nonsmooth Nonconvex Optimization
by: Xiao, Nachuan, et al.
Published: (2024) -
LoCo: Low-Bit Communication Adaptor for Large-scale Model Training
by: Xie, Xingyu, et al.
Published: (2024) -
Stochastic Bregman Subgradient Methods for Nonsmooth Nonconvex Optimization Problems
by: Ding, Kuangyu, et al.
Published: (2024)