Learning Rate Schedules in the Presence of Distribution Shift
Fuente:
arXiv
Saved in:
| Main Authors: | Fahrbach, Matthew, Javanmard, Adel, Mirrokni, Vahab, Worah, Pratik |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PriorBoost: An Adaptive Algorithm for Learning from Aggregate Responses
by: Javanmard, Adel, et al.
Published: (2024)
by: Javanmard, Adel, et al.
Published: (2024)
Less is More: Convergence Benefits of Fewer Data Weight Updates over Longer Horizon
by: Das, Rudrajit, et al.
Published: (2026)
by: Das, Rudrajit, et al.
Published: (2026)
Optimistic Rates for Learning from Label Proportions
by: Li, Gene, et al.
Published: (2024)
by: Li, Gene, et al.
Published: (2024)
Understanding the Role of Training Data in Test-Time Scaling
by: Javanmard, Adel, et al.
Published: (2025)
by: Javanmard, Adel, et al.
Published: (2025)
Theoretical Perspectives on Data Quality and Synergistic Effects in Pre- and Post-Training Reasoning Models
by: Javanmard, Adel, et al.
Published: (2026)
by: Javanmard, Adel, et al.
Published: (2026)
Sampling and Loss Weights in Multi-Domain Training
by: Salmani, Mahdi, et al.
Published: (2025)
by: Salmani, Mahdi, et al.
Published: (2025)
Improving the Variance of Differentially Private Randomized Experiments through Clustering
by: Javanmard, Adel, et al.
Published: (2023)
by: Javanmard, Adel, et al.
Published: (2023)
Learning Optimal Classification Trees Robust to Distribution Shifts
by: Justin, Nathan, et al.
Published: (2023)
by: Justin, Nathan, et al.
Published: (2023)
Optimal Growth Schedules for Batch Size and Learning Rate in SGD that Reduce SFO Complexity
by: Umeda, Hikaru, et al.
Published: (2025)
by: Umeda, Hikaru, et al.
Published: (2025)
Self-Boost via Optimal Retraining: An Analysis via Approximate Message Passing
by: Javanmard, Adel, et al.
Published: (2025)
by: Javanmard, Adel, et al.
Published: (2025)
The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training
by: Schaipp, Fabian, et al.
Published: (2025)
by: Schaipp, Fabian, et al.
Published: (2025)
Boosting Column Generation with Graph Neural Networks for Joint Rider Trip Planning and Crew Shift Scheduling
by: Lu, Jiawei, et al.
Published: (2024)
by: Lu, Jiawei, et al.
Published: (2024)
FedSEA: Achieving Benefit of Parallelization in Federated Online Learning
by: Sahu, Harekrushna, et al.
Published: (2026)
by: Sahu, Harekrushna, et al.
Published: (2026)
Adaptive Batch Size and Learning Rate Scheduler for Stochastic Gradient Descent Based on Minimization of Stochastic First-order Oracle Complexity
by: Umeda, Hikaru, et al.
Published: (2025)
by: Umeda, Hikaru, et al.
Published: (2025)
Mixed-feature Logistic Regression Robust to Distribution Shifts
by: Sun, Qingshi, et al.
Published: (2025)
by: Sun, Qingshi, et al.
Published: (2025)
Load Balancing with Network Latencies via Distributed Gradient Descent
by: Balseiro, Santiago R., et al.
Published: (2025)
by: Balseiro, Santiago R., et al.
Published: (2025)
Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling
by: Meterez, Alexandru, et al.
Published: (2025)
by: Meterez, Alexandru, et al.
Published: (2025)
Anytime Pretraining: Horizon-Free Learning-Rate Schedules with Weight Averaging
by: Meterez, Alexandru, et al.
Published: (2026)
by: Meterez, Alexandru, et al.
Published: (2026)
Heat Death of Generative Models in Closed-Loop Learning
by: Marchi, Matteo, et al.
Published: (2024)
by: Marchi, Matteo, et al.
Published: (2024)
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
Learning from Aggregate responses: Instance Level versus Bag Level Loss Functions
by: Javanmard, Adel, et al.
Published: (2024)
by: Javanmard, Adel, et al.
Published: (2024)
Improved Learning Rates for Stochastic Optimization
by: Li, Shaojie, et al.
Published: (2021)
by: Li, Shaojie, et al.
Published: (2021)
Riemannian coordinate descent algorithms on matrix manifolds
by: Han, Andi, et al.
Published: (2024)
by: Han, Andi, et al.
Published: (2024)
Riemannian Optimization for Hadamard Products of Low-Rank Matrices
by: Jawanpuria, Pratik, et al.
Published: (2026)
by: Jawanpuria, Pratik, et al.
Published: (2026)
Mitigating Covariate Shift in Misspecified Regression with Applications to Reinforcement Learning
by: Amortila, Philip, et al.
Published: (2024)
by: Amortila, Philip, et al.
Published: (2024)
MMD-Regularized Unbalanced Optimal Transport
by: Manupriya, Piyushi, et al.
Published: (2020)
by: Manupriya, Piyushi, et al.
Published: (2020)
Online Learning-guided Learning Rate Adaptation via Gradient Alignment
by: Jiang, Ruichen, et al.
Published: (2025)
by: Jiang, Ruichen, et al.
Published: (2025)
The Marginal Value of Momentum for Small Learning Rate SGD
by: Wang, Runzhe, et al.
Published: (2023)
by: Wang, Runzhe, et al.
Published: (2023)
PROMISE: Preconditioned Stochastic Optimization Methods by Incorporating Scalable Curvature Estimates
by: Frangella, Zachary, et al.
Published: (2023)
by: Frangella, Zachary, et al.
Published: (2023)
SketchySGD: Reliable Stochastic Optimization via Randomized Curvature Estimates
by: Frangella, Zachary, et al.
Published: (2022)
by: Frangella, Zachary, et al.
Published: (2022)
A Framework for Bilevel Optimization on Riemannian Manifolds
by: Han, Andi, et al.
Published: (2024)
by: Han, Andi, et al.
Published: (2024)
Learning Rate Annealing Improves Tuning Robustness in Stochastic Optimization
by: Attia, Amit, et al.
Published: (2025)
by: Attia, Amit, et al.
Published: (2025)
Learning-Rate-Free Stochastic Optimization over Riemannian Manifolds
by: Dodd, Daniel, et al.
Published: (2024)
by: Dodd, Daniel, et al.
Published: (2024)
Provable and Practical Online Learning Rate Adaptation with Hypergradient Descent
by: Chu, Ya-Chi, et al.
Published: (2025)
by: Chu, Ya-Chi, et al.
Published: (2025)
Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence
by: Liu, Yuxing, et al.
Published: (2025)
by: Liu, Yuxing, et al.
Published: (2025)
Design and Scheduling of an AI-based Queueing System
by: Lee, Jiung, et al.
Published: (2024)
by: Lee, Jiung, et al.
Published: (2024)
Distributionally-Robust Learning to Optimize
by: Ranjan, Vinit, et al.
Published: (2026)
by: Ranjan, Vinit, et al.
Published: (2026)
Towards Fast Rates for Federated and Multi-Task Reinforcement Learning
by: Zhu, Feng, et al.
Published: (2024)
by: Zhu, Feng, et al.
Published: (2024)
Corruption-Tolerant Asynchronous Q-Learning with Near-Optimal Rates
by: Maity, Sreejeet, et al.
Published: (2025)
by: Maity, Sreejeet, et al.
Published: (2025)
Understanding Outer Optimizers in Local SGD: Learning Rates, Momentum, and Acceleration
by: Khaled, Ahmed, et al.
Published: (2025)
by: Khaled, Ahmed, et al.
Published: (2025)
Similar Items
-
PriorBoost: An Adaptive Algorithm for Learning from Aggregate Responses
by: Javanmard, Adel, et al.
Published: (2024) -
Less is More: Convergence Benefits of Fewer Data Weight Updates over Longer Horizon
by: Das, Rudrajit, et al.
Published: (2026) -
Optimistic Rates for Learning from Label Proportions
by: Li, Gene, et al.
Published: (2024) -
Understanding the Role of Training Data in Test-Time Scaling
by: Javanmard, Adel, et al.
Published: (2025) -
Theoretical Perspectives on Data Quality and Synergistic Effects in Pre- and Post-Training Reasoning Models
by: Javanmard, Adel, et al.
Published: (2026)