Anytime Pretraining: Horizon-Free Learning-Rate Schedules with Weight Averaging
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Meterez, Alexandru, Nair, Pranav Ajit, Morwani, Depen, Pehlevan, Cengiz, Kakade, Sham |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling
par: Meterez, Alexandru, et autres
Publié: (2025)
par: Meterez, Alexandru, et autres
Publié: (2025)
A Simplified Analysis of SGD for Linear Regression with Weight Averaging
par: Meterez, Alexandru, et autres
Publié: (2025)
par: Meterez, Alexandru, et autres
Publié: (2025)
How Does Critical Batch Size Scale in Pre-training?
par: Zhang, Hanlin, et autres
Publié: (2024)
par: Zhang, Hanlin, et autres
Publié: (2024)
Connections between Schedule-Free Optimizers, AdEMAMix, and Accelerated SGD Variants
par: Morwani, Depen, et autres
Publié: (2025)
par: Morwani, Depen, et autres
Publié: (2025)
Anytime Training with Schedule-Free Spectral Optimization
par: Apte, Anuj, et autres
Publié: (2026)
par: Apte, Anuj, et autres
Publié: (2026)
A New Perspective on Shampoo's Preconditioner
par: Morwani, Depen, et autres
Publié: (2024)
par: Morwani, Depen, et autres
Publié: (2024)
The Potential of Second-Order Optimization for LLMs: A Study with Full Gauss-Newton
par: Abreu, Natalie, et autres
Publié: (2025)
par: Abreu, Natalie, et autres
Publié: (2025)
Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
par: Zhao, Rosie, et autres
Publié: (2025)
par: Zhao, Rosie, et autres
Publié: (2025)
Deconstructing What Makes a Good Optimizer for Language Models
par: Zhao, Rosie, et autres
Publié: (2024)
par: Zhao, Rosie, et autres
Publié: (2024)
Learning-Guided Rolling Horizon Optimization for Long-Horizon Flexible Job-Shop Scheduling
par: Li, Sirui, et autres
Publié: (2025)
par: Li, Sirui, et autres
Publié: (2025)
The Recurrent Transformer: Greater Effective Depth and Efficient Decoding
par: Oncescu, Costin-Andrei, et autres
Publié: (2026)
par: Oncescu, Costin-Andrei, et autres
Publié: (2026)
Convex Relaxation for Solving Large-Margin Classifiers in Hyperbolic Space
par: Yang, Sheng, et autres
Publié: (2024)
par: Yang, Sheng, et autres
Publié: (2024)
Through the River: Understanding the Benefit of Schedule-Free Methods for Language Model Training
par: Song, Minhak, et autres
Publié: (2025)
par: Song, Minhak, et autres
Publié: (2025)
SOAP: Improving and Stabilizing Shampoo using Adam
par: Vyas, Nikhil, et autres
Publié: (2024)
par: Vyas, Nikhil, et autres
Publié: (2024)
The Road Less Scheduled
par: Defazio, Aaron, et autres
Publié: (2024)
par: Defazio, Aaron, et autres
Publié: (2024)
ART for Diffusion Sampling: A Reinforcement Learning Approach to Timestep Schedule
par: Huang, Yilie, et autres
Publié: (2026)
par: Huang, Yilie, et autres
Publié: (2026)
Integrated Offline and Online Learning to Solve a Large Class of Scheduling Problems
par: Liu, Anbang, et autres
Publié: (2025)
par: Liu, Anbang, et autres
Publié: (2025)
Infinite-Horizon Reach-Avoid Zero-Sum Games via Deep Reinforcement Learning
par: Li, Jingqi, et autres
Publié: (2022)
par: Li, Jingqi, et autres
Publié: (2022)
Memory-Efficient LLM Pretraining via Minimalist Optimizer Design
par: Glentis, Athanasios, et autres
Publié: (2025)
par: Glentis, Athanasios, et autres
Publié: (2025)
Hindsight-Guided Momentum (HGM) Optimizer: An Approach to Adaptive Learning Rate
par: Sarkar, Krisanu
Publié: (2025)
par: Sarkar, Krisanu
Publié: (2025)
Solving Integrated Process Planning and Scheduling Problem via Graph Neural Network Based Deep Reinforcement Learning
par: Li, Hongpei, et autres
Publié: (2024)
par: Li, Hongpei, et autres
Publié: (2024)
Online Scheduling for LLM Inference with KV Cache Constraints
par: Jaillet, Patrick, et autres
Publié: (2025)
par: Jaillet, Patrick, et autres
Publié: (2025)
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less
par: Liu, Yuxing, et autres
Publié: (2026)
par: Liu, Yuxing, et autres
Publié: (2026)
Matching the Statistical Query Lower Bound for $k$-Sparse Parity Problems with Sign Stochastic Gradient Descent
par: Kou, Yiwen, et autres
Publié: (2024)
par: Kou, Yiwen, et autres
Publié: (2024)
Revisiting LQR Control from the Perspective of Receding-Horizon Policy Gradient
par: Zhang, Xiangyuan, et autres
Publié: (2023)
par: Zhang, Xiangyuan, et autres
Publié: (2023)
Graph Neural Networks for the Offline Nanosatellite Task Scheduling Problem
par: Pacheco, Bruno Machado, et autres
Publié: (2023)
par: Pacheco, Bruno Machado, et autres
Publié: (2023)
Neural Combinatorial Optimization for Stochastic Flexible Job Shop Scheduling Problems
par: Smit, Igor G., et autres
Publié: (2024)
par: Smit, Igor G., et autres
Publié: (2024)
Optimization and Generalization Guarantees for Weight Normalization
par: Cisneros-Velarde, Pedro, et autres
Publié: (2024)
par: Cisneros-Velarde, Pedro, et autres
Publié: (2024)
Lagrangian Index Policy for Restless Bandits with Average Reward
par: Avrachenkov, Konstantin, et autres
Publié: (2024)
par: Avrachenkov, Konstantin, et autres
Publié: (2024)
Open Problem: Anytime Convergence Rate of Gradient Descent
par: Kornowski, Guy, et autres
Publié: (2024)
par: Kornowski, Guy, et autres
Publié: (2024)
Weighted Low-rank Approximation via Stochastic Gradient Descent on Manifolds
par: Xu, Conglong, et autres
Publié: (2025)
par: Xu, Conglong, et autres
Publié: (2025)
Approximate and Weighted Data Reconstruction Attack in Federated Learning
par: Song, Yongcun, et autres
Publié: (2023)
par: Song, Yongcun, et autres
Publié: (2023)
Prodigy: An Expeditiously Adaptive Parameter-Free Learner
par: Mishchenko, Konstantin, et autres
Publié: (2023)
par: Mishchenko, Konstantin, et autres
Publié: (2023)
Beyond Minimax Rates in Group Distributionally Robust Optimization via a Novel Notion of Sparsity
par: Nguyen, Quan, et autres
Publié: (2024)
par: Nguyen, Quan, et autres
Publié: (2024)
Achieving Tighter Finite-Time Rates for Heterogeneous Federated Stochastic Approximation under Markovian Sampling
par: Zhu, Feng, et autres
Publié: (2025)
par: Zhu, Feng, et autres
Publié: (2025)
Gating is Weighting: Understanding Gated Linear Attention through In-context Learning
par: Li, Yingcong, et autres
Publié: (2025)
par: Li, Yingcong, et autres
Publié: (2025)
BAGEL: Projection-Free Algorithm for Adversarially Constrained Online Convex Optimization
par: Lu, Yiyang, et autres
Publié: (2025)
par: Lu, Yiyang, et autres
Publié: (2025)
PyCSP3-Scheduling: A Scheduling Extension for PyCSP3
par: Afifi, Sohaib
Publié: (2026)
par: Afifi, Sohaib
Publié: (2026)
Deep Reinforcement Learning for Flexible Job Shop Scheduling with Random Job Arrivals
par: Tang, Yu, et autres
Publié: (2026)
par: Tang, Yu, et autres
Publié: (2026)
Kernel-Free Universum Quadratic Surface Twin Support Vector Machines for Imbalanced Data
par: Moosaei, Hossein, et autres
Publié: (2024)
par: Moosaei, Hossein, et autres
Publié: (2024)
Documents similaires
-
Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling
par: Meterez, Alexandru, et autres
Publié: (2025) -
A Simplified Analysis of SGD for Linear Regression with Weight Averaging
par: Meterez, Alexandru, et autres
Publié: (2025) -
How Does Critical Batch Size Scale in Pre-training?
par: Zhang, Hanlin, et autres
Publié: (2024) -
Connections between Schedule-Free Optimizers, AdEMAMix, and Accelerated SGD Variants
par: Morwani, Depen, et autres
Publié: (2025) -
Anytime Training with Schedule-Free Spectral Optimization
par: Apte, Anuj, et autres
Publié: (2026)