Anytime Pretraining: Horizon-Free Learning-Rate Schedules with Weight Averaging
Fuente:
arXiv
Saved in:
| Main Authors: | Meterez, Alexandru, Nair, Pranav Ajit, Morwani, Depen, Pehlevan, Cengiz, Kakade, Sham |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling
by: Meterez, Alexandru, et al.
Published: (2025)
by: Meterez, Alexandru, et al.
Published: (2025)
A Simplified Analysis of SGD for Linear Regression with Weight Averaging
by: Meterez, Alexandru, et al.
Published: (2025)
by: Meterez, Alexandru, et al.
Published: (2025)
How Does Critical Batch Size Scale in Pre-training?
by: Zhang, Hanlin, et al.
Published: (2024)
by: Zhang, Hanlin, et al.
Published: (2024)
Connections between Schedule-Free Optimizers, AdEMAMix, and Accelerated SGD Variants
by: Morwani, Depen, et al.
Published: (2025)
by: Morwani, Depen, et al.
Published: (2025)
Anytime Training with Schedule-Free Spectral Optimization
by: Apte, Anuj, et al.
Published: (2026)
by: Apte, Anuj, et al.
Published: (2026)
A New Perspective on Shampoo's Preconditioner
by: Morwani, Depen, et al.
Published: (2024)
by: Morwani, Depen, et al.
Published: (2024)
The Potential of Second-Order Optimization for LLMs: A Study with Full Gauss-Newton
by: Abreu, Natalie, et al.
Published: (2025)
by: Abreu, Natalie, et al.
Published: (2025)
Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
by: Zhao, Rosie, et al.
Published: (2025)
by: Zhao, Rosie, et al.
Published: (2025)
Deconstructing What Makes a Good Optimizer for Language Models
by: Zhao, Rosie, et al.
Published: (2024)
by: Zhao, Rosie, et al.
Published: (2024)
Learning-Guided Rolling Horizon Optimization for Long-Horizon Flexible Job-Shop Scheduling
by: Li, Sirui, et al.
Published: (2025)
by: Li, Sirui, et al.
Published: (2025)
The Recurrent Transformer: Greater Effective Depth and Efficient Decoding
by: Oncescu, Costin-Andrei, et al.
Published: (2026)
by: Oncescu, Costin-Andrei, et al.
Published: (2026)
Convex Relaxation for Solving Large-Margin Classifiers in Hyperbolic Space
by: Yang, Sheng, et al.
Published: (2024)
by: Yang, Sheng, et al.
Published: (2024)
Through the River: Understanding the Benefit of Schedule-Free Methods for Language Model Training
by: Song, Minhak, et al.
Published: (2025)
by: Song, Minhak, et al.
Published: (2025)
SOAP: Improving and Stabilizing Shampoo using Adam
by: Vyas, Nikhil, et al.
Published: (2024)
by: Vyas, Nikhil, et al.
Published: (2024)
The Road Less Scheduled
by: Defazio, Aaron, et al.
Published: (2024)
by: Defazio, Aaron, et al.
Published: (2024)
ART for Diffusion Sampling: A Reinforcement Learning Approach to Timestep Schedule
by: Huang, Yilie, et al.
Published: (2026)
by: Huang, Yilie, et al.
Published: (2026)
Integrated Offline and Online Learning to Solve a Large Class of Scheduling Problems
by: Liu, Anbang, et al.
Published: (2025)
by: Liu, Anbang, et al.
Published: (2025)
Infinite-Horizon Reach-Avoid Zero-Sum Games via Deep Reinforcement Learning
by: Li, Jingqi, et al.
Published: (2022)
by: Li, Jingqi, et al.
Published: (2022)
Memory-Efficient LLM Pretraining via Minimalist Optimizer Design
by: Glentis, Athanasios, et al.
Published: (2025)
by: Glentis, Athanasios, et al.
Published: (2025)
Hindsight-Guided Momentum (HGM) Optimizer: An Approach to Adaptive Learning Rate
by: Sarkar, Krisanu
Published: (2025)
by: Sarkar, Krisanu
Published: (2025)
Solving Integrated Process Planning and Scheduling Problem via Graph Neural Network Based Deep Reinforcement Learning
by: Li, Hongpei, et al.
Published: (2024)
by: Li, Hongpei, et al.
Published: (2024)
Online Scheduling for LLM Inference with KV Cache Constraints
by: Jaillet, Patrick, et al.
Published: (2025)
by: Jaillet, Patrick, et al.
Published: (2025)
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less
by: Liu, Yuxing, et al.
Published: (2026)
by: Liu, Yuxing, et al.
Published: (2026)
Matching the Statistical Query Lower Bound for $k$-Sparse Parity Problems with Sign Stochastic Gradient Descent
by: Kou, Yiwen, et al.
Published: (2024)
by: Kou, Yiwen, et al.
Published: (2024)
Revisiting LQR Control from the Perspective of Receding-Horizon Policy Gradient
by: Zhang, Xiangyuan, et al.
Published: (2023)
by: Zhang, Xiangyuan, et al.
Published: (2023)
Graph Neural Networks for the Offline Nanosatellite Task Scheduling Problem
by: Pacheco, Bruno Machado, et al.
Published: (2023)
by: Pacheco, Bruno Machado, et al.
Published: (2023)
Neural Combinatorial Optimization for Stochastic Flexible Job Shop Scheduling Problems
by: Smit, Igor G., et al.
Published: (2024)
by: Smit, Igor G., et al.
Published: (2024)
Optimization and Generalization Guarantees for Weight Normalization
by: Cisneros-Velarde, Pedro, et al.
Published: (2024)
by: Cisneros-Velarde, Pedro, et al.
Published: (2024)
Lagrangian Index Policy for Restless Bandits with Average Reward
by: Avrachenkov, Konstantin, et al.
Published: (2024)
by: Avrachenkov, Konstantin, et al.
Published: (2024)
Open Problem: Anytime Convergence Rate of Gradient Descent
by: Kornowski, Guy, et al.
Published: (2024)
by: Kornowski, Guy, et al.
Published: (2024)
Weighted Low-rank Approximation via Stochastic Gradient Descent on Manifolds
by: Xu, Conglong, et al.
Published: (2025)
by: Xu, Conglong, et al.
Published: (2025)
Approximate and Weighted Data Reconstruction Attack in Federated Learning
by: Song, Yongcun, et al.
Published: (2023)
by: Song, Yongcun, et al.
Published: (2023)
Prodigy: An Expeditiously Adaptive Parameter-Free Learner
by: Mishchenko, Konstantin, et al.
Published: (2023)
by: Mishchenko, Konstantin, et al.
Published: (2023)
Beyond Minimax Rates in Group Distributionally Robust Optimization via a Novel Notion of Sparsity
by: Nguyen, Quan, et al.
Published: (2024)
by: Nguyen, Quan, et al.
Published: (2024)
Achieving Tighter Finite-Time Rates for Heterogeneous Federated Stochastic Approximation under Markovian Sampling
by: Zhu, Feng, et al.
Published: (2025)
by: Zhu, Feng, et al.
Published: (2025)
Gating is Weighting: Understanding Gated Linear Attention through In-context Learning
by: Li, Yingcong, et al.
Published: (2025)
by: Li, Yingcong, et al.
Published: (2025)
BAGEL: Projection-Free Algorithm for Adversarially Constrained Online Convex Optimization
by: Lu, Yiyang, et al.
Published: (2025)
by: Lu, Yiyang, et al.
Published: (2025)
PyCSP3-Scheduling: A Scheduling Extension for PyCSP3
by: Afifi, Sohaib
Published: (2026)
by: Afifi, Sohaib
Published: (2026)
Deep Reinforcement Learning for Flexible Job Shop Scheduling with Random Job Arrivals
by: Tang, Yu, et al.
Published: (2026)
by: Tang, Yu, et al.
Published: (2026)
Kernel-Free Universum Quadratic Surface Twin Support Vector Machines for Imbalanced Data
by: Moosaei, Hossein, et al.
Published: (2024)
by: Moosaei, Hossein, et al.
Published: (2024)
Similar Items
-
Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling
by: Meterez, Alexandru, et al.
Published: (2025) -
A Simplified Analysis of SGD for Linear Regression with Weight Averaging
by: Meterez, Alexandru, et al.
Published: (2025) -
How Does Critical Batch Size Scale in Pre-training?
by: Zhang, Hanlin, et al.
Published: (2024) -
Connections between Schedule-Free Optimizers, AdEMAMix, and Accelerated SGD Variants
by: Morwani, Depen, et al.
Published: (2025) -
Anytime Training with Schedule-Free Spectral Optimization
by: Apte, Anuj, et al.
Published: (2026)