Predictable Scale: Part I, Step Law -- Optimal Hyperparameter Scaling Law in Large Language Model Pretraining
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Houyi, Zheng, Wenzhen, Wang, Qiufeng, Zhang, Hanshan, Wang, Zili, Xuyang, Shijie, Fan, Yuantao, Ding, Zhenyu, Wang, Haoying, Ding, Ning, Zhou, Shuigeng, Zhang, Xiangyu, Jiang, Daxin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Predictable Scale: Part II, Farseer: A Refined Scaling Law in Large Language Models
di: Li, Houyi, et al.
Pubblicazione: (2025)
di: Li, Houyi, et al.
Pubblicazione: (2025)
The Data Efficiency Frontier of Financial Foundation Models: Scaling Laws from Continued Pretraining
di: Ponnock, Jesse
Pubblicazione: (2025)
di: Ponnock, Jesse
Pubblicazione: (2025)
Getting ViT in Shape: Scaling Laws for Compute-Optimal Model Design
di: Alabdulmohsin, Ibrahim, et al.
Pubblicazione: (2023)
di: Alabdulmohsin, Ibrahim, et al.
Pubblicazione: (2023)
Scaling Laws for State Dynamics in Large Language Models
di: Li, Jacob X, et al.
Pubblicazione: (2025)
di: Li, Jacob X, et al.
Pubblicazione: (2025)
Heaven & Hell II: Scale Laws and Robustness in One-Step Heaven-Hell Consensus
di: Aghanya, Nnamdi Daniel, et al.
Pubblicazione: (2025)
di: Aghanya, Nnamdi Daniel, et al.
Pubblicazione: (2025)
Scaling Laws for Neural Material Models
di: Trikha, Akshay, et al.
Pubblicazione: (2025)
di: Trikha, Akshay, et al.
Pubblicazione: (2025)
Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource
di: Li, Houyi, et al.
Pubblicazione: (2025)
di: Li, Houyi, et al.
Pubblicazione: (2025)
Hardware Co-Design Scaling Laws via Roofline Modelling for On-Device LLMs
di: Sun, Luoyang, et al.
Pubblicazione: (2026)
di: Sun, Luoyang, et al.
Pubblicazione: (2026)
Closing the Curvature Gap: Full Transformer Hessians and Their Implications for Scaling Laws
di: Petrov, Egor, et al.
Pubblicazione: (2025)
di: Petrov, Egor, et al.
Pubblicazione: (2025)
Towards Greater Leverage: Scaling Laws for Efficient Mixture-of-Experts Language Models
di: Tian, Changxin, et al.
Pubblicazione: (2025)
di: Tian, Changxin, et al.
Pubblicazione: (2025)
ProbeScale: Probing Analysis to Optimize Neural Scaling Laws for Efficient Small Language Model Inference
di: Das, Sourav
Pubblicazione: (2026)
di: Das, Sourav
Pubblicazione: (2026)
Law-Strength Frontiers and a No-Free-Lunch Result for Law-Seeking Reinforcement Learning on Volatility Law Manifolds
di: Zhang, Jian'an
Pubblicazione: (2025)
di: Zhang, Jian'an
Pubblicazione: (2025)
Step-Audio-R1 Technical Report
di: Tian, Fei, et al.
Pubblicazione: (2025)
di: Tian, Fei, et al.
Pubblicazione: (2025)
LLM-Driven Large-Scale Spectrum Access
di: Yang, Ning, et al.
Pubblicazione: (2026)
di: Yang, Ning, et al.
Pubblicazione: (2026)
Singularity-aware Optimization via Randomized Geometric Probing: Towards Stable Non-smooth Optimization
di: Xu, Ruoran, et al.
Pubblicazione: (2026)
di: Xu, Ruoran, et al.
Pubblicazione: (2026)
Near-Optimal Wafer-Scale Reduce
di: Luczynski, Piotr, et al.
Pubblicazione: (2024)
di: Luczynski, Piotr, et al.
Pubblicazione: (2024)
Scaling Laws for Associative Memories
di: Cabannes, Vivien, et al.
Pubblicazione: (2023)
di: Cabannes, Vivien, et al.
Pubblicazione: (2023)
Multi-matrix Factorization Attention
di: Hu, Jingcheng, et al.
Pubblicazione: (2024)
di: Hu, Jingcheng, et al.
Pubblicazione: (2024)
Scaling Laws for Forgetting When Fine-Tuning Large Language Models
di: Kalajdzievski, Damjan
Pubblicazione: (2024)
di: Kalajdzievski, Damjan
Pubblicazione: (2024)
Optimizing Sequential Recommendation Models with Scaling Laws and Approximate Entropy
di: Shen, Tingjia, et al.
Pubblicazione: (2024)
di: Shen, Tingjia, et al.
Pubblicazione: (2024)
Murphys Laws of AI Alignment: Why the Gap Always Wins
di: Gaikwad, Madhava
Pubblicazione: (2025)
di: Gaikwad, Madhava
Pubblicazione: (2025)
Testing the Efficacy of Hyperparameter Optimization Algorithms in Short-Term Load Forecasting
di: Hakyemez, Tugrul Cabir, et al.
Pubblicazione: (2024)
di: Hakyemez, Tugrul Cabir, et al.
Pubblicazione: (2024)
Spectra: Surprising Effectiveness of Pretraining Ternary Language Models at Scale
di: Kaushal, Ayush, et al.
Pubblicazione: (2024)
di: Kaushal, Ayush, et al.
Pubblicazione: (2024)
LLM Architecture, Scaling Laws, and Economics: A Quick Summary
di: Press, William H.
Pubblicazione: (2025)
di: Press, William H.
Pubblicazione: (2025)
Evidence for a Functional Proximity Law in Multilayer Networks
di: Ivanov, Vladi
Pubblicazione: (2026)
di: Ivanov, Vladi
Pubblicazione: (2026)
Hilbert-Geo: Solving Solid Geometric Problems by Neural-Symbolic Reasoning
di: Xu, Ruoran, et al.
Pubblicazione: (2026)
di: Xu, Ruoran, et al.
Pubblicazione: (2026)
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning
di: Mircea, Andrei, et al.
Pubblicazione: (2025)
di: Mircea, Andrei, et al.
Pubblicazione: (2025)
DISC: Dynamic Decomposition Improves LLM Inference Scaling
di: Light, Jonathan, et al.
Pubblicazione: (2025)
di: Light, Jonathan, et al.
Pubblicazione: (2025)
By Their Fruits You Will Know Them: Comparing Formalizations of Law by the Decisions They Encode
di: Vernie, Julius, et al.
Pubblicazione: (2026)
di: Vernie, Julius, et al.
Pubblicazione: (2026)
CaMeRL: Collision-Aware and Memory-Enhanced Reinforcement Learning for UAV Navigation in Multi-Scale Obstacle Environments
di: Hong, Hong, et al.
Pubblicazione: (2026)
di: Hong, Hong, et al.
Pubblicazione: (2026)
The Reward Function and the Least Cost Principle for Gravitation and other Laws of Physics
di: Moreno-Bote, Rubén
Pubblicazione: (2026)
di: Moreno-Bote, Rubén
Pubblicazione: (2026)
Deterministic Minimum Steiner Cut in Maximum Flow Time
di: Ding, Matthew, et al.
Pubblicazione: (2023)
di: Ding, Matthew, et al.
Pubblicazione: (2023)
Corpus Considerations for Annotator Modeling and Scaling
di: Sarumi, Olufunke O., et al.
Pubblicazione: (2024)
di: Sarumi, Olufunke O., et al.
Pubblicazione: (2024)
Large-Scale LiDAR-Inertial Dataset for Degradation-Robust High-Precision Mapping
di: Jin, Xiaofeng, et al.
Pubblicazione: (2025)
di: Jin, Xiaofeng, et al.
Pubblicazione: (2025)
Scaling Laws for Code: A More Data-Hungry Regime
di: Luo, Xianzhen, et al.
Pubblicazione: (2025)
di: Luo, Xianzhen, et al.
Pubblicazione: (2025)
FORGE: An LLM-driven Framework for Large-Scale Smart Contract Vulnerability Dataset Construction
di: Chen, Jiachi, et al.
Pubblicazione: (2025)
di: Chen, Jiachi, et al.
Pubblicazione: (2025)
ContractBench: Can LLM Agents Preserve Observation Contracts?
di: Wang, Jicheng, et al.
Pubblicazione: (2026)
di: Wang, Jicheng, et al.
Pubblicazione: (2026)
Scaling Laws in the Tiny Regime: How Small Models Change Their Mistakes
di: Alnemari, Mohammed, et al.
Pubblicazione: (2026)
di: Alnemari, Mohammed, et al.
Pubblicazione: (2026)
DRFormer: Multi-Scale Transformer Utilizing Diverse Receptive Fields for Long Time-Series Forecasting
di: Ding, Ruixin, et al.
Pubblicazione: (2024)
di: Ding, Ruixin, et al.
Pubblicazione: (2024)
StepScorer: Accelerating Reinforcement Learning with Step-wise Scoring and Psychological Regret Modeling
di: Xu, Zhe
Pubblicazione: (2026)
di: Xu, Zhe
Pubblicazione: (2026)
Documenti analoghi
-
Predictable Scale: Part II, Farseer: A Refined Scaling Law in Large Language Models
di: Li, Houyi, et al.
Pubblicazione: (2025) -
The Data Efficiency Frontier of Financial Foundation Models: Scaling Laws from Continued Pretraining
di: Ponnock, Jesse
Pubblicazione: (2025) -
Getting ViT in Shape: Scaling Laws for Compute-Optimal Model Design
di: Alabdulmohsin, Ibrahim, et al.
Pubblicazione: (2023) -
Scaling Laws for State Dynamics in Large Language Models
di: Li, Jacob X, et al.
Pubblicazione: (2025) -
Heaven & Hell II: Scale Laws and Robustness in One-Step Heaven-Hell Consensus
di: Aghanya, Nnamdi Daniel, et al.
Pubblicazione: (2025)