Through the River: Understanding the Benefit of Schedule-Free Methods for Language Model Training
Fuente:
arXiv
Saved in:
| Main Authors: | Song, Minhak, Baek, Beomhan, Ahn, Kwangjun, Yun, Chulhee |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Implicit Bias of Per-sample Adam on Separable Data: Departure from the Full-batch Regime
by: Baek, Beomhan, et al.
Published: (2025)
by: Baek, Beomhan, et al.
Published: (2025)
Does SGD really happen in tiny subspaces?
by: Song, Minhak, et al.
Published: (2024)
by: Song, Minhak, et al.
Published: (2024)
Linear attention is (maybe) all you need (to understand transformer optimization)
by: Ahn, Kwangjun, et al.
Published: (2023)
by: Ahn, Kwangjun, et al.
Published: (2023)
How to escape sharp minima with random perturbations
by: Ahn, Kwangjun, et al.
Published: (2023)
by: Ahn, Kwangjun, et al.
Published: (2023)
Scaling Laws of SignSGD in Linear Regression: When Does It Outperform SGD?
by: Kim, Jihwan, et al.
Published: (2026)
by: Kim, Jihwan, et al.
Published: (2026)
Anytime Training with Schedule-Free Spectral Optimization
by: Apte, Anuj, et al.
Published: (2026)
by: Apte, Anuj, et al.
Published: (2026)
General framework for online-to-nonconvex conversion: Schedule-free SGD is also effective for nonconvex optimization
by: Ahn, Kwangjun, et al.
Published: (2024)
by: Ahn, Kwangjun, et al.
Published: (2024)
Learning Weakly Communicating Average-Reward CMDPs: Strong Duality and Improved Regret
by: Yu, Kihyun, et al.
Published: (2026)
by: Yu, Kihyun, et al.
Published: (2026)
Fundamental Benefit of Alternating Updates in Minimax Optimization
by: Lee, Jaewook, et al.
Published: (2024)
by: Lee, Jaewook, et al.
Published: (2024)
Dion: Distributed Orthonormalized Updates
by: Ahn, Kwangjun, et al.
Published: (2025)
by: Ahn, Kwangjun, et al.
Published: (2025)
Adam with model exponential moving average is effective for nonconvex optimization
by: Ahn, Kwangjun, et al.
Published: (2024)
by: Ahn, Kwangjun, et al.
Published: (2024)
Understanding Adam Optimizer via Online Learning of Updates: Adam is FTRL in Disguise
by: Ahn, Kwangjun, et al.
Published: (2024)
by: Ahn, Kwangjun, et al.
Published: (2024)
Understanding Sampler Stochasticity in Training Diffusion Models for RLHF
by: Sheng, Jiayuan, et al.
Published: (2025)
by: Sheng, Jiayuan, et al.
Published: (2025)
Anytime Pretraining: Horizon-Free Learning-Rate Schedules with Weight Averaging
by: Meterez, Alexandru, et al.
Published: (2026)
by: Meterez, Alexandru, et al.
Published: (2026)
Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling
by: Meterez, Alexandru, et al.
Published: (2025)
by: Meterez, Alexandru, et al.
Published: (2025)
The Sharpness Disparity Principle in Transformers for Accelerating Language Model Pre-Training
by: Wang, Jinbo, et al.
Published: (2025)
by: Wang, Jinbo, et al.
Published: (2025)
Preconditioning Benefits of Spectral Orthogonalization in Muon
by: Ma, Jianhao, et al.
Published: (2026)
by: Ma, Jianhao, et al.
Published: (2026)
The Road Less Scheduled
by: Defazio, Aaron, et al.
Published: (2024)
by: Defazio, Aaron, et al.
Published: (2024)
Unsupervised Training of Diffusion Models for Feasible Solution Generation in Neural Combinatorial Optimization
by: Hong, Seong-Hyun, et al.
Published: (2024)
by: Hong, Seong-Hyun, et al.
Published: (2024)
Provable Acceleration of Nesterov's Accelerated Gradient Method over Heavy Ball Method in Training Over-Parameterized Neural Networks
by: Liu, Xin, et al.
Published: (2022)
by: Liu, Xin, et al.
Published: (2022)
StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models
by: Yu, Dingzhi, et al.
Published: (2026)
by: Yu, Dingzhi, et al.
Published: (2026)
Incremental Gradient Descent with Small Epoch Counts is Surprisingly Slow on Ill-Conditioned Problems
by: Kim, Yujun, et al.
Published: (2025)
by: Kim, Yujun, et al.
Published: (2025)
Convergence and Implicit Bias of Gradient Descent on Continual Linear Classification
by: Jung, Hyunji, et al.
Published: (2025)
by: Jung, Hyunji, et al.
Published: (2025)
Stochastic Extragradient with Flip-Flop Shuffling & Anchoring: Provable Improvements
by: Chae, Jiseok, et al.
Published: (2024)
by: Chae, Jiseok, et al.
Published: (2024)
Online Scheduling for LLM Inference with KV Cache Constraints
by: Jaillet, Patrick, et al.
Published: (2025)
by: Jaillet, Patrick, et al.
Published: (2025)
Graph Neural Networks for the Offline Nanosatellite Task Scheduling Problem
by: Pacheco, Bruno Machado, et al.
Published: (2023)
by: Pacheco, Bruno Machado, et al.
Published: (2023)
Understanding Sharpness Dynamics in NN Training with a Minimalist Example: The Effects of Dataset Difficulty, Depth, Stochasticity, and More
by: Yoo, Geonhui, et al.
Published: (2025)
by: Yoo, Geonhui, et al.
Published: (2025)
Neural Combinatorial Optimization for Stochastic Flexible Job Shop Scheduling Problems
by: Smit, Igor G., et al.
Published: (2024)
by: Smit, Igor G., et al.
Published: (2024)
Enhancing Stochastic Gradient Descent: A Unified Framework and Novel Acceleration Methods for Faster Convergence
by: Deng, Yichuan, et al.
Published: (2024)
by: Deng, Yichuan, et al.
Published: (2024)
Integrated Offline and Online Learning to Solve a Large Class of Scheduling Problems
by: Liu, Anbang, et al.
Published: (2025)
by: Liu, Anbang, et al.
Published: (2025)
Learning-Guided Rolling Horizon Optimization for Long-Horizon Flexible Job-Shop Scheduling
by: Li, Sirui, et al.
Published: (2025)
by: Li, Sirui, et al.
Published: (2025)
Isotropic Curvature Model for Understanding Deep Learning Optimization: Is Gradient Orthogonalization Optimal?
by: Su, Weijie
Published: (2025)
by: Su, Weijie
Published: (2025)
ART for Diffusion Sampling: A Reinforcement Learning Approach to Timestep Schedule
by: Huang, Yilie, et al.
Published: (2026)
by: Huang, Yilie, et al.
Published: (2026)
Solving Integrated Process Planning and Scheduling Problem via Graph Neural Network Based Deep Reinforcement Learning
by: Li, Hongpei, et al.
Published: (2024)
by: Li, Hongpei, et al.
Published: (2024)
Understanding Fixed Predictions via Confined Regions
by: Lawless, Connor, et al.
Published: (2025)
by: Lawless, Connor, et al.
Published: (2025)
Understanding Optimization in Deep Learning with Central Flows
by: Cohen, Jeremy M., et al.
Published: (2024)
by: Cohen, Jeremy M., et al.
Published: (2024)
A Theoretical Framework for Auxiliary-Loss-Free Load Balancing of Sparse Mixture-of-Experts in Large-Scale AI Models
by: Han, X. Y., et al.
Published: (2025)
by: Han, X. Y., et al.
Published: (2025)
Training Infinitely Deep and Wide Transformers
by: Barboni, Raphaël, et al.
Published: (2026)
by: Barboni, Raphaël, et al.
Published: (2026)
Prodigy: An Expeditiously Adaptive Parameter-Free Learner
by: Mishchenko, Konstantin, et al.
Published: (2023)
by: Mishchenko, Konstantin, et al.
Published: (2023)
Accelerating RLHF Training with Reward Variance Increase
by: Yang, Zonglin, et al.
Published: (2025)
by: Yang, Zonglin, et al.
Published: (2025)
Similar Items
-
Implicit Bias of Per-sample Adam on Separable Data: Departure from the Full-batch Regime
by: Baek, Beomhan, et al.
Published: (2025) -
Does SGD really happen in tiny subspaces?
by: Song, Minhak, et al.
Published: (2024) -
Linear attention is (maybe) all you need (to understand transformer optimization)
by: Ahn, Kwangjun, et al.
Published: (2023) -
How to escape sharp minima with random perturbations
by: Ahn, Kwangjun, et al.
Published: (2023) -
Scaling Laws of SignSGD in Linear Regression: When Does It Outperform SGD?
by: Kim, Jihwan, et al.
Published: (2026)