Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective
Fuente:
arXiv
Salvato in:
| Autori principali: | Wen, Kaiyue, Li, Zhiyuan, Wang, Jason, Hall, David, Liang, Percy, Ma, Tengyu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Fantastic Pretraining Optimizers and Where to Find Them
di: Wen, Kaiyue, et al.
Pubblicazione: (2025)
di: Wen, Kaiyue, et al.
Pubblicazione: (2025)
Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training
di: Liu, Hong, et al.
Pubblicazione: (2023)
di: Liu, Hong, et al.
Pubblicazione: (2023)
Optimal Learning-Rate Schedules under Functional Scaling Laws: Power Decay and Warmup-Stable-Decay
di: Li, Binghui, et al.
Pubblicazione: (2026)
di: Li, Binghui, et al.
Pubblicazione: (2026)
Training Dynamics of the Cooldown Stage in Warmup-Stable-Decay Learning Rate Scheduler
di: Dremov, Aleksandr, et al.
Pubblicazione: (2025)
di: Dremov, Aleksandr, et al.
Pubblicazione: (2025)
Configuration-to-Performance Scaling Law with Neural Ansatz
di: Zhang, Huaqing, et al.
Pubblicazione: (2026)
di: Zhang, Huaqing, et al.
Pubblicazione: (2026)
Divide-and-Conquer CoT: RL for Reducing Latency via Parallel Reasoning
di: Mahankali, Arvind, et al.
Pubblicazione: (2026)
di: Mahankali, Arvind, et al.
Pubblicazione: (2026)
Universal Dynamics of Warmup Stable Decay: understanding WSD beyond Transformers
di: Belloni, Annalisa, et al.
Pubblicazione: (2026)
di: Belloni, Annalisa, et al.
Pubblicazione: (2026)
Power-Law Decay Loss for Large Language Model Finetuning: A Theory Perspective
di: Shao, Jintian
Pubblicazione: (2025)
di: Shao, Jintian
Pubblicazione: (2025)
A Multi-Power Law for Loss Curve Prediction Across Learning Rate Schedules
di: Luo, Kairong, et al.
Pubblicazione: (2025)
di: Luo, Kairong, et al.
Pubblicazione: (2025)
Analyzing & Reducing the Need for Learning Rate Warmup in GPT Training
di: Kosson, Atli, et al.
Pubblicazione: (2024)
di: Kosson, Atli, et al.
Pubblicazione: (2024)
Taming Transformer Without Using Learning Rate Warmup
di: Qi, Xianbiao, et al.
Pubblicazione: (2025)
di: Qi, Xianbiao, et al.
Pubblicazione: (2025)
RNNs are not Transformers (Yet): The Key Bottleneck on In-context Retrieval
di: Wen, Kaiyue, et al.
Pubblicazione: (2024)
di: Wen, Kaiyue, et al.
Pubblicazione: (2024)
Scaling Self-Play with Self-Guidance
di: Bailey, Luke, et al.
Pubblicazione: (2026)
di: Bailey, Luke, et al.
Pubblicazione: (2026)
Why Warmup the Learning Rate? Underlying Mechanisms and Improvements
di: Kalra, Dayal Singh, et al.
Pubblicazione: (2024)
di: Kalra, Dayal Singh, et al.
Pubblicazione: (2024)
Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence
di: Liu, Yuxing, et al.
Pubblicazione: (2025)
di: Liu, Yuxing, et al.
Pubblicazione: (2025)
Understanding the Generalization Benefits of Late Learning Rate Decay
di: Ren, Yinuo, et al.
Pubblicazione: (2024)
di: Ren, Yinuo, et al.
Pubblicazione: (2024)
Replaying pre-training data improves fine-tuning
di: Kotha, Suhas, et al.
Pubblicazione: (2026)
di: Kotha, Suhas, et al.
Pubblicazione: (2026)
Self-Verified Distillation: Your Language Model Is Secretly Its Own Synthetic Data Pipeline
di: Lee, Tony, et al.
Pubblicazione: (2026)
di: Lee, Tony, et al.
Pubblicazione: (2026)
AdaLRS: Loss-Guided Adaptive Learning Rate Search for Efficient Foundation Model Pretraining
di: Dong, Hongyuan, et al.
Pubblicazione: (2025)
di: Dong, Hongyuan, et al.
Pubblicazione: (2025)
Analyzing Consumer Reviews for Understanding Drivers of Hotels Ratings: An Indian Perspective
di: Dasgupta, Subhasis, et al.
Pubblicazione: (2024)
di: Dasgupta, Subhasis, et al.
Pubblicazione: (2024)
Pre-training LLM without Learning Rate Decay Enhances Supervised Fine-Tuning
di: Yano, Kazuki, et al.
Pubblicazione: (2026)
di: Yano, Kazuki, et al.
Pubblicazione: (2026)
How Learning Rate Decay Wastes Your Best Data in Curriculum-Based LLM Pretraining
di: Luo, Kairong, et al.
Pubblicazione: (2025)
di: Luo, Kairong, et al.
Pubblicazione: (2025)
Non-Asymptotic Length Generalization
di: Chen, Thomas, et al.
Pubblicazione: (2025)
di: Chen, Thomas, et al.
Pubblicazione: (2025)
Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models
di: Qiu, Zihan, et al.
Pubblicazione: (2025)
di: Qiu, Zihan, et al.
Pubblicazione: (2025)
From Sparse Dependence to Sparse Attention: Unveiling How Chain-of-Thought Enhances Transformer Sample Efficiency
di: Wen, Kaiyue, et al.
Pubblicazione: (2024)
di: Wen, Kaiyue, et al.
Pubblicazione: (2024)
Understanding Emergent Abilities of Language Models from the Loss Perspective
di: Du, Zhengxiao, et al.
Pubblicazione: (2024)
di: Du, Zhengxiao, et al.
Pubblicazione: (2024)
Pseudo-Formalization for Automatic Proof Verification
di: Barkallah, Slim, et al.
Pubblicazione: (2026)
di: Barkallah, Slim, et al.
Pubblicazione: (2026)
Linguistic Calibration of Long-Form Generations
di: Band, Neil, et al.
Pubblicazione: (2024)
di: Band, Neil, et al.
Pubblicazione: (2024)
Loss Landscape Degeneracy and Stagewise Development in Transformers
di: Hoogland, Jesse, et al.
Pubblicazione: (2024)
di: Hoogland, Jesse, et al.
Pubblicazione: (2024)
LK Losses: Direct Acceptance Rate Optimization for Speculative Decoding
di: Samarin, Alexander, et al.
Pubblicazione: (2026)
di: Samarin, Alexander, et al.
Pubblicazione: (2026)
Independence Tests for Language Models
di: Zhu, Sally, et al.
Pubblicazione: (2025)
di: Zhu, Sally, et al.
Pubblicazione: (2025)
Convex Dominance in Deep Learning I: A Scaling Law of Loss and Learning Rate
di: Bu, Zhiqi, et al.
Pubblicazione: (2026)
di: Bu, Zhiqi, et al.
Pubblicazione: (2026)
Beyond a Single Extractor: Re-thinking HTML-to-Text Extraction for LLM Pretraining
di: Li, Jeffrey, et al.
Pubblicazione: (2026)
di: Li, Jeffrey, et al.
Pubblicazione: (2026)
WSM: Decay-Free Learning Rate Schedule via Checkpoint Merging for LLM Pre-training
di: Tian, Changxin, et al.
Pubblicazione: (2025)
di: Tian, Changxin, et al.
Pubblicazione: (2025)
On the Entropy Calibration of Language Models
di: Cao, Steven, et al.
Pubblicazione: (2025)
di: Cao, Steven, et al.
Pubblicazione: (2025)
Cross-Model Comparative Loss for Enhancing Neuronal Utility in Language Understanding
di: Zhu, Yunchang, et al.
Pubblicazione: (2023)
di: Zhu, Yunchang, et al.
Pubblicazione: (2023)
Large Language Models as Tool Makers
di: Cai, Tianle, et al.
Pubblicazione: (2023)
di: Cai, Tianle, et al.
Pubblicazione: (2023)
PEFT-Arena: Understanding Parameter-Efficient Finetuning from a Stability-Plasticity Perspective
di: Huang, Yangyi, et al.
Pubblicazione: (2026)
di: Huang, Yangyi, et al.
Pubblicazione: (2026)
More Edits, More Stable: Understanding the Lifelong Normalization in Sequential Model Editing
di: Ma, Xin, et al.
Pubblicazione: (2026)
di: Ma, Xin, et al.
Pubblicazione: (2026)
On the Learnability of Watermarks for Language Models
di: Gu, Chenchen, et al.
Pubblicazione: (2023)
di: Gu, Chenchen, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Fantastic Pretraining Optimizers and Where to Find Them
di: Wen, Kaiyue, et al.
Pubblicazione: (2025) -
Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training
di: Liu, Hong, et al.
Pubblicazione: (2023) -
Optimal Learning-Rate Schedules under Functional Scaling Laws: Power Decay and Warmup-Stable-Decay
di: Li, Binghui, et al.
Pubblicazione: (2026) -
Training Dynamics of the Cooldown Stage in Warmup-Stable-Decay Learning Rate Scheduler
di: Dremov, Aleksandr, et al.
Pubblicazione: (2025) -
Configuration-to-Performance Scaling Law with Neural Ansatz
di: Zhang, Huaqing, et al.
Pubblicazione: (2026)