Power Lines: Scaling Laws for Weight Decay and Batch Size in LLM Pre-training
Fuente:
arXiv
Saved in:
| Main Authors: | Bergsma, Shane, Dey, Nolan, Gosal, Gurpreet, Gray, Gavia, Soboleva, Daria, Hestness, Joel |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Straight to Zero: Why Linearly Decaying the Learning Rate to Zero Works Best for LLMs
by: Bergsma, Shane, et al.
Published: (2025)
by: Bergsma, Shane, et al.
Published: (2025)
Scaling with Collapse: Efficient and Predictable Training of LLM Families
by: Bergsma, Shane, et al.
Published: (2025)
by: Bergsma, Shane, et al.
Published: (2025)
Predicting Training Re-evaluation Curves Enables Effective Data Curriculums for LLMs
by: Bergsma, Shane, et al.
Published: (2025)
by: Bergsma, Shane, et al.
Published: (2025)
Normalization Layer Per-Example Gradients are Sufficient to Predict Gradient Noise Scale in Transformers
by: Gray, Gavia, et al.
Published: (2024)
by: Gray, Gavia, et al.
Published: (2024)
Sparse maximal update parameterization: A holistic approach to sparse training dynamics
by: Dey, Nolan, et al.
Published: (2024)
by: Dey, Nolan, et al.
Published: (2024)
PTPP-Aware Adaptation Scaling Laws: Predicting Domain-Adaptation Performance at Unseen Pre-Training Budgets
by: Goffinet, Etienne, et al.
Published: (2025)
by: Goffinet, Etienne, et al.
Published: (2025)
SlimPajama-DC: Understanding Data Combinations for LLM Training
by: Shen, Zhiqiang, et al.
Published: (2023)
by: Shen, Zhiqiang, et al.
Published: (2023)
Don't be lazy: CompleteP enables compute-efficient deep transformers
by: Dey, Nolan, et al.
Published: (2025)
by: Dey, Nolan, et al.
Published: (2025)
Scaling LLM Pre-training with Vocabulary Curriculum
by: Yu, Fangyuan
Published: (2025)
by: Yu, Fangyuan
Published: (2025)
How to Set the Batch Size for Large-Scale Pre-training?
by: Zhou, Yunhua, et al.
Published: (2026)
by: Zhou, Yunhua, et al.
Published: (2026)
Demystifying Synthetic Data in LLM Pre-training: A Systematic Study of Scaling Laws, Benefits, and Pitfalls
by: Kang, Feiyang, et al.
Published: (2025)
by: Kang, Feiyang, et al.
Published: (2025)
How Does Critical Batch Size Scale in Pre-training?
by: Zhang, Hanlin, et al.
Published: (2024)
by: Zhang, Hanlin, et al.
Published: (2024)
MediSwift: Efficient Sparse Pre-trained Biomedical Language Models
by: Thangarasa, Vithursan, et al.
Published: (2024)
by: Thangarasa, Vithursan, et al.
Published: (2024)
How to inject knowledge efficiently? Knowledge Infusion Scaling Law for Pre-training Large Language Models
by: Lv, Kangtao, et al.
Published: (2025)
by: Lv, Kangtao, et al.
Published: (2025)
Batched Contextual Reinforcement: A Task-Scaling Law for Efficient Reasoning
by: Yang, Bangji, et al.
Published: (2026)
by: Yang, Bangji, et al.
Published: (2026)
Scaling Law for Language Models Training Considering Batch Size
by: Shuai, Xian, et al.
Published: (2024)
by: Shuai, Xian, et al.
Published: (2024)
Power Scheduler: A Batch Size and Token Number Agnostic Learning Rate Scheduler
by: Shen, Yikang, et al.
Published: (2024)
by: Shen, Yikang, et al.
Published: (2024)
Bilingual Adaptation of Monolingual Foundation Models
by: Gosal, Gurpreet, et al.
Published: (2024)
by: Gosal, Gurpreet, et al.
Published: (2024)
The Scaling Laws of Skills in LLM Agent Systems
by: Chen, Charles, et al.
Published: (2026)
by: Chen, Charles, et al.
Published: (2026)
Batch Speculative Decoding Done Right
by: Zhang, Ranran Haoran, et al.
Published: (2025)
by: Zhang, Ranran Haoran, et al.
Published: (2025)
LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws
by: Mayilvahanan, Prasanna, et al.
Published: (2025)
by: Mayilvahanan, Prasanna, et al.
Published: (2025)
Bridging the Fairness Gap: Enhancing Pre-trained Models with LLM-Generated Sentences
by: Yu, Liu, et al.
Published: (2025)
by: Yu, Liu, et al.
Published: (2025)
AlphaDecay: Module-wise Weight Decay for Heavy-Tailed Balancing in LLMs
by: He, Di, et al.
Published: (2025)
by: He, Di, et al.
Published: (2025)
Crystal: Illuminating LLM Abilities on Language and Code
by: Tao, Tianhua, et al.
Published: (2024)
by: Tao, Tianhua, et al.
Published: (2024)
PLDR-LLM: Large Language Model from Power Law Decoder Representations
by: Gokden, Burc
Published: (2024)
by: Gokden, Burc
Published: (2024)
GQA-μP: The maximal parameterization update for grouped query attention
by: Chickering, Kyle R., et al.
Published: (2026)
by: Chickering, Kyle R., et al.
Published: (2026)
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs
by: Bi, Baolong, et al.
Published: (2025)
by: Bi, Baolong, et al.
Published: (2025)
Calibrating Pre-trained Language Classifiers on LLM-generated Noisy Labels via Iterative Refinement
by: Ye, Liqin, et al.
Published: (2025)
by: Ye, Liqin, et al.
Published: (2025)
Evaluating LLM-based Approaches to Legal Citation Prediction: Domain-specific Pre-training, Fine-tuning, or RAG? A Benchmark and an Australian Law Case Study
by: Han, Jiuzhou, et al.
Published: (2024)
by: Han, Jiuzhou, et al.
Published: (2024)
BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching
by: Zheng, Zhen, et al.
Published: (2024)
by: Zheng, Zhen, et al.
Published: (2024)
Crown, Frame, Reverse: Layer-Wise Scaling Variants for LLM Pre-Training
by: Baroian, Andrei, et al.
Published: (2025)
by: Baroian, Andrei, et al.
Published: (2025)
SmolTulu: Higher Learning Rate to Batch Size Ratios Can Lead to Better Reasoning in SLMs
by: Alrashed, Sultan
Published: (2024)
by: Alrashed, Sultan
Published: (2024)
Variance Control via Weight Rescaling in LLM Pre-training
by: Owen, Louis, et al.
Published: (2025)
by: Owen, Louis, et al.
Published: (2025)
Programming Every Example: Lifting Pre-training Data Quality Like Experts at Scale
by: Zhou, Fan, et al.
Published: (2024)
by: Zhou, Fan, et al.
Published: (2024)
Scaling Smart: Accelerating Large Language Model Pre-training with Small Model Initialization
by: Samragh, Mohammad, et al.
Published: (2024)
by: Samragh, Mohammad, et al.
Published: (2024)
NuNER: Entity Recognition Encoder Pre-training via LLM-Annotated Data
by: Bogdanov, Sergei, et al.
Published: (2024)
by: Bogdanov, Sergei, et al.
Published: (2024)
Adaptive Layer-skipping in Pre-trained LLMs
by: Luo, Xuan, et al.
Published: (2025)
by: Luo, Xuan, et al.
Published: (2025)
Exploring the Benefit of Activation Sparsity in Pre-training
by: Zhang, Zhengyan, et al.
Published: (2024)
by: Zhang, Zhengyan, et al.
Published: (2024)
PDR: A Plug-and-Play Positional Decay Framework for LLM Pre-training Data Detection
by: Liu, Jinhan, et al.
Published: (2026)
by: Liu, Jinhan, et al.
Published: (2026)
How do Scaling Laws Apply to Knowledge Graph Engineering Tasks? The Impact of Model Size on Large Language Model Performance
by: Heim, Desiree, et al.
Published: (2025)
by: Heim, Desiree, et al.
Published: (2025)
Similar Items
-
Straight to Zero: Why Linearly Decaying the Learning Rate to Zero Works Best for LLMs
by: Bergsma, Shane, et al.
Published: (2025) -
Scaling with Collapse: Efficient and Predictable Training of LLM Families
by: Bergsma, Shane, et al.
Published: (2025) -
Predicting Training Re-evaluation Curves Enables Effective Data Curriculums for LLMs
by: Bergsma, Shane, et al.
Published: (2025) -
Normalization Layer Per-Example Gradients are Sufficient to Predict Gradient Noise Scale in Transformers
by: Gray, Gavia, et al.
Published: (2024) -
Sparse maximal update parameterization: A holistic approach to sparse training dynamics
by: Dey, Nolan, et al.
Published: (2024)