Surge Phenomenon in Optimal Learning Rate and Batch Size Scaling
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Shuaipeng, Zhao, Penghao, Zhang, Hailin, Sun, Xingwu, Wu, Hao, Jiao, Dian, Wang, Weiyan, Liu, Chengjun, Fang, Zheng, Xue, Jinbao, Tao, Yangyu, Cui, Bin, Wang, Di |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MEMO: Fine-grained Tensor Management For Ultra-long Context LLM Training
by: Zhao, Pinxue, et al.
Published: (2024)
by: Zhao, Pinxue, et al.
Published: (2024)
BeamVQ: Aligning Space-Time Forecasting Model via Self-training on Physics-aware Metrics
by: Wu, Hao, et al.
Published: (2024)
by: Wu, Hao, et al.
Published: (2024)
BeamVQ: Beam Search with Vector Quantization to Mitigate Data Scarcity in Physical Spatiotemporal Forecasting
by: Wang, Weiyan, et al.
Published: (2025)
by: Wang, Weiyan, et al.
Published: (2025)
Scaling Laws for Floating Point Quantization Training
by: Sun, Xingwu, et al.
Published: (2025)
by: Sun, Xingwu, et al.
Published: (2025)
Learning from History: A Retrieval-Augmented Framework for Spatiotemporal Prediction
by: Jia, Hao, et al.
Published: (2025)
by: Jia, Hao, et al.
Published: (2025)
Exploiting Student Parallelism for Efficient GPU Inference of BERT-like Models in Online Services
by: Wang, Weiyan, et al.
Published: (2024)
by: Wang, Weiyan, et al.
Published: (2024)
Hydraulis: Balancing Large Transformer Model Training via Co-designing Parallel Strategies and Data Assignment
by: Li, Haoyang, et al.
Published: (2024)
by: Li, Haoyang, et al.
Published: (2024)
Time Transfer: On Optimal Learning Rate and Batch Size In The Infinite Data Limit
by: Filatov, Oleg, et al.
Published: (2024)
by: Filatov, Oleg, et al.
Published: (2024)
Self-Distillation for Multi-Token Prediction
by: Zhao, Guoliang, et al.
Published: (2026)
by: Zhao, Guoliang, et al.
Published: (2026)
On Optimal Batch Size in Coded Computing
by: Saha, Swapnil, et al.
Published: (2025)
by: Saha, Swapnil, et al.
Published: (2025)
Fast Catch-Up, Late Switching: Optimal Batch Size Scheduling via Functional Scaling Laws
by: Wang, Jinbo, et al.
Published: (2026)
by: Wang, Jinbo, et al.
Published: (2026)
Optimal Growth Schedules for Batch Size and Learning Rate in SGD that Reduce SFO Complexity
by: Umeda, Hikaru, et al.
Published: (2025)
by: Umeda, Hikaru, et al.
Published: (2025)
Strategic Joining and Optimal Pricing in a Single‐Server Batch Arrival Queue With Different Information of Batch Size
by: Kaili Li, et al.
Published: (2025)
by: Kaili Li, et al.
Published: (2025)
More Expressive Attention with Negative Weights
by: Lv, Ang, et al.
Published: (2024)
by: Lv, Ang, et al.
Published: (2024)
Scaling Law for Language Models Training Considering Batch Size
by: Shuai, Xian, et al.
Published: (2024)
by: Shuai, Xian, et al.
Published: (2024)
Optimal Batch-Size Control for Low-Latency Federated Learning with Device Heterogeneity
by: Yang, Huiling, et al.
Published: (2025)
by: Yang, Huiling, et al.
Published: (2025)
Convergence of Riemannian Stochastic Gradient Descents: Varying Batch Sizes And Nonstandard Batch Forming
by: Wu, Hao
Published: (2026)
by: Wu, Hao
Published: (2026)
Towards a Comprehensive Scaling Law of Mixture-of-Experts
by: Zhao, Guoliang, et al.
Published: (2025)
by: Zhao, Guoliang, et al.
Published: (2025)
Experimental Analysis of Large-scale Learnable Vector Storage Compression
by: Zhang, Hailin, et al.
Published: (2023)
by: Zhang, Hailin, et al.
Published: (2023)
Grazing-sliding bifurcations in planar $\mathbb{Z}_2$-symmetric Filippov systems
by: Chen, Xingwu, et al.
Published: (2025)
by: Chen, Xingwu, et al.
Published: (2025)
Bifurcations of grazing loops of arbitrary tangent multiplicity in piecewise-smooth systems
by: Chen, Xingwu, et al.
Published: (2026)
by: Chen, Xingwu, et al.
Published: (2026)
Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling
by: Meterez, Alexandru, et al.
Published: (2025)
by: Meterez, Alexandru, et al.
Published: (2025)
Truth Forest: Toward Multi-Scale Truthfulness in Large Language Models through Intervention without Tuning
by: Chen, Zhongzhi, et al.
Published: (2023)
by: Chen, Zhongzhi, et al.
Published: (2023)
How to Set the Batch Size for Large-Scale Pre-training?
by: Zhou, Yunhua, et al.
Published: (2026)
by: Zhou, Yunhua, et al.
Published: (2026)
How Does Critical Batch Size Scale in Pre-training?
by: Zhang, Hanlin, et al.
Published: (2024)
by: Zhang, Hanlin, et al.
Published: (2024)
Increasing Both Batch Size and Learning Rate Accelerates Stochastic Gradient Descent
by: Umeda, Hikaru, et al.
Published: (2024)
by: Umeda, Hikaru, et al.
Published: (2024)
AdaBatchGrad: Combining Adaptive Batch Size and Adaptive Step Size
by: Ostroukhov, Petr, et al.
Published: (2024)
by: Ostroukhov, Petr, et al.
Published: (2024)
Classifications and bifurcations of tangent points and their loops of planar piecewise-smooth systems
by: Fang, Zhihao, et al.
Published: (2024)
by: Fang, Zhihao, et al.
Published: (2024)
Bifurcations and explicit unfoldings of grazing loops connecting one high multiplicity tangent point
by: Fang, Zhihao, et al.
Published: (2024)
by: Fang, Zhihao, et al.
Published: (2024)
Gap Phenomenon for Yamabe Type Problems of $M^m\times T^{n-m}$
by: Wang, Fang, et al.
Published: (2026)
by: Wang, Fang, et al.
Published: (2026)
Power Scheduler: A Batch Size and Token Number Agnostic Learning Rate Scheduler
by: Shen, Yikang, et al.
Published: (2024)
by: Shen, Yikang, et al.
Published: (2024)
Retrieval-Augmented Generation for AI-Generated Content: A Survey
by: Zhao, Penghao, et al.
Published: (2024)
by: Zhao, Penghao, et al.
Published: (2024)
Adaptive Batch Sizes Using Non-Euclidean Gradient Noise Scales for Stochastic Sign and Spectral Descent
by: Naganuma, Hiroki, et al.
Published: (2026)
by: Naganuma, Hiroki, et al.
Published: (2026)
Breaking the Memory Barrier: Near Infinite Batch Size Scaling for Contrastive Loss
by: Cheng, Zesen, et al.
Published: (2024)
by: Cheng, Zesen, et al.
Published: (2024)
Deep Fuzzy Optimization for Batch-Size and Nearest Neighbors in Optimal Robot Motion Planning
by: Zhang, Liding, et al.
Published: (2025)
by: Zhang, Liding, et al.
Published: (2025)
Accelerating SGDM via Learning Rate and Batch Size Schedules: A Lyapunov-Based Analysis
by: Kondo, Yuichi, et al.
Published: (2025)
by: Kondo, Yuichi, et al.
Published: (2025)
Convergence of Sharpness-Aware Minimization Algorithms using Increasing Batch Size and Decaying Learning Rate
by: Harada, Hinata, et al.
Published: (2024)
by: Harada, Hinata, et al.
Published: (2024)
Recent Progress on the Surface and Bulk Modification of Carbon Nitrides for Photocatalytic Carbon Dioxide Reduction
by: Siqing Ma, et al.
Published: (2024)
by: Siqing Ma, et al.
Published: (2024)
Proof of a conjecture of Green and Liebeck on codes in symmetric groups
by: Fang, Teng, et al.
Published: (2025)
by: Fang, Teng, et al.
Published: (2025)
TransMamba: A Sequence-Level Hybrid Transformer-Mamba Language Model
by: Li, Yixing, et al.
Published: (2025)
by: Li, Yixing, et al.
Published: (2025)
Similar Items
-
MEMO: Fine-grained Tensor Management For Ultra-long Context LLM Training
by: Zhao, Pinxue, et al.
Published: (2024) -
BeamVQ: Aligning Space-Time Forecasting Model via Self-training on Physics-aware Metrics
by: Wu, Hao, et al.
Published: (2024) -
BeamVQ: Beam Search with Vector Quantization to Mitigate Data Scarcity in Physical Spatiotemporal Forecasting
by: Wang, Weiyan, et al.
Published: (2025) -
Scaling Laws for Floating Point Quantization Training
by: Sun, Xingwu, et al.
Published: (2025) -
Learning from History: A Retrieval-Augmented Framework for Spatiotemporal Prediction
by: Jia, Hao, et al.
Published: (2025)