Predictable Scale: Part I, Step Law -- Optimal Hyperparameter Scaling Law in Large Language Model Pretraining
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Li, Houyi, Zheng, Wenzhen, Wang, Qiufeng, Zhang, Hanshan, Wang, Zili, Xuyang, Shijie, Fan, Yuantao, Ding, Zhenyu, Wang, Haoying, Ding, Ning, Zhou, Shuigeng, Zhang, Xiangyu, Jiang, Daxin |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Predictable Scale: Part II, Farseer: A Refined Scaling Law in Large Language Models
par: Li, Houyi, et autres
Publié: (2025)
par: Li, Houyi, et autres
Publié: (2025)
The Data Efficiency Frontier of Financial Foundation Models: Scaling Laws from Continued Pretraining
par: Ponnock, Jesse
Publié: (2025)
par: Ponnock, Jesse
Publié: (2025)
Getting ViT in Shape: Scaling Laws for Compute-Optimal Model Design
par: Alabdulmohsin, Ibrahim, et autres
Publié: (2023)
par: Alabdulmohsin, Ibrahim, et autres
Publié: (2023)
Scaling Laws for State Dynamics in Large Language Models
par: Li, Jacob X, et autres
Publié: (2025)
par: Li, Jacob X, et autres
Publié: (2025)
Heaven & Hell II: Scale Laws and Robustness in One-Step Heaven-Hell Consensus
par: Aghanya, Nnamdi Daniel, et autres
Publié: (2025)
par: Aghanya, Nnamdi Daniel, et autres
Publié: (2025)
Scaling Laws for Neural Material Models
par: Trikha, Akshay, et autres
Publié: (2025)
par: Trikha, Akshay, et autres
Publié: (2025)
Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource
par: Li, Houyi, et autres
Publié: (2025)
par: Li, Houyi, et autres
Publié: (2025)
Hardware Co-Design Scaling Laws via Roofline Modelling for On-Device LLMs
par: Sun, Luoyang, et autres
Publié: (2026)
par: Sun, Luoyang, et autres
Publié: (2026)
Closing the Curvature Gap: Full Transformer Hessians and Their Implications for Scaling Laws
par: Petrov, Egor, et autres
Publié: (2025)
par: Petrov, Egor, et autres
Publié: (2025)
Towards Greater Leverage: Scaling Laws for Efficient Mixture-of-Experts Language Models
par: Tian, Changxin, et autres
Publié: (2025)
par: Tian, Changxin, et autres
Publié: (2025)
ProbeScale: Probing Analysis to Optimize Neural Scaling Laws for Efficient Small Language Model Inference
par: Das, Sourav
Publié: (2026)
par: Das, Sourav
Publié: (2026)
Law-Strength Frontiers and a No-Free-Lunch Result for Law-Seeking Reinforcement Learning on Volatility Law Manifolds
par: Zhang, Jian'an
Publié: (2025)
par: Zhang, Jian'an
Publié: (2025)
Step-Audio-R1 Technical Report
par: Tian, Fei, et autres
Publié: (2025)
par: Tian, Fei, et autres
Publié: (2025)
LLM-Driven Large-Scale Spectrum Access
par: Yang, Ning, et autres
Publié: (2026)
par: Yang, Ning, et autres
Publié: (2026)
Singularity-aware Optimization via Randomized Geometric Probing: Towards Stable Non-smooth Optimization
par: Xu, Ruoran, et autres
Publié: (2026)
par: Xu, Ruoran, et autres
Publié: (2026)
Near-Optimal Wafer-Scale Reduce
par: Luczynski, Piotr, et autres
Publié: (2024)
par: Luczynski, Piotr, et autres
Publié: (2024)
Scaling Laws for Associative Memories
par: Cabannes, Vivien, et autres
Publié: (2023)
par: Cabannes, Vivien, et autres
Publié: (2023)
Multi-matrix Factorization Attention
par: Hu, Jingcheng, et autres
Publié: (2024)
par: Hu, Jingcheng, et autres
Publié: (2024)
Scaling Laws for Forgetting When Fine-Tuning Large Language Models
par: Kalajdzievski, Damjan
Publié: (2024)
par: Kalajdzievski, Damjan
Publié: (2024)
Optimizing Sequential Recommendation Models with Scaling Laws and Approximate Entropy
par: Shen, Tingjia, et autres
Publié: (2024)
par: Shen, Tingjia, et autres
Publié: (2024)
Murphys Laws of AI Alignment: Why the Gap Always Wins
par: Gaikwad, Madhava
Publié: (2025)
par: Gaikwad, Madhava
Publié: (2025)
Testing the Efficacy of Hyperparameter Optimization Algorithms in Short-Term Load Forecasting
par: Hakyemez, Tugrul Cabir, et autres
Publié: (2024)
par: Hakyemez, Tugrul Cabir, et autres
Publié: (2024)
Spectra: Surprising Effectiveness of Pretraining Ternary Language Models at Scale
par: Kaushal, Ayush, et autres
Publié: (2024)
par: Kaushal, Ayush, et autres
Publié: (2024)
LLM Architecture, Scaling Laws, and Economics: A Quick Summary
par: Press, William H.
Publié: (2025)
par: Press, William H.
Publié: (2025)
Evidence for a Functional Proximity Law in Multilayer Networks
par: Ivanov, Vladi
Publié: (2026)
par: Ivanov, Vladi
Publié: (2026)
Hilbert-Geo: Solving Solid Geometric Problems by Neural-Symbolic Reasoning
par: Xu, Ruoran, et autres
Publié: (2026)
par: Xu, Ruoran, et autres
Publié: (2026)
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning
par: Mircea, Andrei, et autres
Publié: (2025)
par: Mircea, Andrei, et autres
Publié: (2025)
DISC: Dynamic Decomposition Improves LLM Inference Scaling
par: Light, Jonathan, et autres
Publié: (2025)
par: Light, Jonathan, et autres
Publié: (2025)
By Their Fruits You Will Know Them: Comparing Formalizations of Law by the Decisions They Encode
par: Vernie, Julius, et autres
Publié: (2026)
par: Vernie, Julius, et autres
Publié: (2026)
CaMeRL: Collision-Aware and Memory-Enhanced Reinforcement Learning for UAV Navigation in Multi-Scale Obstacle Environments
par: Hong, Hong, et autres
Publié: (2026)
par: Hong, Hong, et autres
Publié: (2026)
The Reward Function and the Least Cost Principle for Gravitation and other Laws of Physics
par: Moreno-Bote, Rubén
Publié: (2026)
par: Moreno-Bote, Rubén
Publié: (2026)
Deterministic Minimum Steiner Cut in Maximum Flow Time
par: Ding, Matthew, et autres
Publié: (2023)
par: Ding, Matthew, et autres
Publié: (2023)
Corpus Considerations for Annotator Modeling and Scaling
par: Sarumi, Olufunke O., et autres
Publié: (2024)
par: Sarumi, Olufunke O., et autres
Publié: (2024)
Large-Scale LiDAR-Inertial Dataset for Degradation-Robust High-Precision Mapping
par: Jin, Xiaofeng, et autres
Publié: (2025)
par: Jin, Xiaofeng, et autres
Publié: (2025)
Scaling Laws for Code: A More Data-Hungry Regime
par: Luo, Xianzhen, et autres
Publié: (2025)
par: Luo, Xianzhen, et autres
Publié: (2025)
FORGE: An LLM-driven Framework for Large-Scale Smart Contract Vulnerability Dataset Construction
par: Chen, Jiachi, et autres
Publié: (2025)
par: Chen, Jiachi, et autres
Publié: (2025)
ContractBench: Can LLM Agents Preserve Observation Contracts?
par: Wang, Jicheng, et autres
Publié: (2026)
par: Wang, Jicheng, et autres
Publié: (2026)
Scaling Laws in the Tiny Regime: How Small Models Change Their Mistakes
par: Alnemari, Mohammed, et autres
Publié: (2026)
par: Alnemari, Mohammed, et autres
Publié: (2026)
DRFormer: Multi-Scale Transformer Utilizing Diverse Receptive Fields for Long Time-Series Forecasting
par: Ding, Ruixin, et autres
Publié: (2024)
par: Ding, Ruixin, et autres
Publié: (2024)
StepScorer: Accelerating Reinforcement Learning with Step-wise Scoring and Psychological Regret Modeling
par: Xu, Zhe
Publié: (2026)
par: Xu, Zhe
Publié: (2026)
Documents similaires
-
Predictable Scale: Part II, Farseer: A Refined Scaling Law in Large Language Models
par: Li, Houyi, et autres
Publié: (2025) -
The Data Efficiency Frontier of Financial Foundation Models: Scaling Laws from Continued Pretraining
par: Ponnock, Jesse
Publié: (2025) -
Getting ViT in Shape: Scaling Laws for Compute-Optimal Model Design
par: Alabdulmohsin, Ibrahim, et autres
Publié: (2023) -
Scaling Laws for State Dynamics in Large Language Models
par: Li, Jacob X, et autres
Publié: (2025) -
Heaven & Hell II: Scale Laws and Robustness in One-Step Heaven-Hell Consensus
par: Aghanya, Nnamdi Daniel, et autres
Publié: (2025)