Unifying Learning Dynamics and Generalization in Transformers Scaling Law
Fuente:
arXiv
Saved in:
| Main Author: | Yang, Chiwun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unlocking the Theory Behind Scaling 1-Bit Neural Networks
by: Daliri, Majid, et al.
Published: (2024)
by: Daliri, Majid, et al.
Published: (2024)
Towards Infinite-Long Prefix in Transformer
by: Liang, Yingyu, et al.
Published: (2024)
by: Liang, Yingyu, et al.
Published: (2024)
How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse
by: Deng, Yichuan, et al.
Published: (2024)
by: Deng, Yichuan, et al.
Published: (2024)
Diversity of Transformer Layers: One Aspect of Parameter Scaling Laws
by: Kamigaito, Hidetaka, et al.
Published: (2025)
by: Kamigaito, Hidetaka, et al.
Published: (2025)
Scaling Law with Learning Rate Annealing
by: Tissue, Howe, et al.
Published: (2024)
by: Tissue, Howe, et al.
Published: (2024)
A Hitchhiker's Guide to Scaling Law Estimation
by: Choshen, Leshem, et al.
Published: (2024)
by: Choshen, Leshem, et al.
Published: (2024)
Distillation Scaling Laws
by: Busbridge, Dan, et al.
Published: (2025)
by: Busbridge, Dan, et al.
Published: (2025)
Scaling Laws for Predicting Downstream Performance in LLMs
by: Chen, Yangyi, et al.
Published: (2024)
by: Chen, Yangyi, et al.
Published: (2024)
What Scales in Cross-Entropy Scaling Law?
by: Yan, Junxi, et al.
Published: (2025)
by: Yan, Junxi, et al.
Published: (2025)
Understanding Scaling Laws with Statistical and Approximation Theory for Transformer Neural Networks on Intrinsically Low-dimensional Data
by: Havrilla, Alex, et al.
Published: (2024)
by: Havrilla, Alex, et al.
Published: (2024)
Batched Contextual Reinforcement: A Task-Scaling Law for Efficient Reasoning
by: Yang, Bangji, et al.
Published: (2026)
by: Yang, Bangji, et al.
Published: (2026)
Towards a Comprehensive Scaling Law of Mixture-of-Experts
by: Zhao, Guoliang, et al.
Published: (2025)
by: Zhao, Guoliang, et al.
Published: (2025)
Scaling Laws for Many-Shot In-Context Learning with Self-Generated Annotations
by: Gu, Zhengyao, et al.
Published: (2025)
by: Gu, Zhengyao, et al.
Published: (2025)
P$^2$ Law: Scaling Law for Post-Training After Model Pruning
by: Chen, Xiaodong, et al.
Published: (2024)
by: Chen, Xiaodong, et al.
Published: (2024)
A Tale of Tails: Model Collapse as a Change of Scaling Laws
by: Dohmatob, Elvis, et al.
Published: (2024)
by: Dohmatob, Elvis, et al.
Published: (2024)
Can Language Models Discover Scaling Laws?
by: Lin, Haowei, et al.
Published: (2025)
by: Lin, Haowei, et al.
Published: (2025)
Exploring Scaling Laws for EHR Foundation Models
by: Zhang, Sheng, et al.
Published: (2025)
by: Zhang, Sheng, et al.
Published: (2025)
Theoretical Foundations of Scaling Law in Familial Models
by: Song, Huan, et al.
Published: (2025)
by: Song, Huan, et al.
Published: (2025)
Scaling Laws for Fine-Grained Mixture of Experts
by: Krajewski, Jakub, et al.
Published: (2024)
by: Krajewski, Jakub, et al.
Published: (2024)
MaskTab: Scalable Masked Tabular Pretraining with Scaling Laws and Distillation for Industrial Classification
by: Zheng, Bo, et al.
Published: (2026)
by: Zheng, Bo, et al.
Published: (2026)
Inverse Reinforcement Learning with Dynamic Reward Scaling for LLM Alignment
by: Cheng, Ruoxi, et al.
Published: (2025)
by: Cheng, Ruoxi, et al.
Published: (2025)
Predicting Task Performance with Context-aware Scaling Laws
by: Montgomery, Kyle, et al.
Published: (2025)
by: Montgomery, Kyle, et al.
Published: (2025)
Relative-Based Scaling Law for Neural Language Models
by: Yue, Baoqing, et al.
Published: (2025)
by: Yue, Baoqing, et al.
Published: (2025)
Loss-to-Loss Prediction: Scaling Laws for All Datasets
by: Brandfonbrener, David, et al.
Published: (2024)
by: Brandfonbrener, David, et al.
Published: (2024)
Observational Scaling Laws and the Predictability of Language Model Performance
by: Ruan, Yangjun, et al.
Published: (2024)
by: Ruan, Yangjun, et al.
Published: (2024)
To Memorize or to Retrieve: Scaling Laws for RAG-Considerate Pretraining
by: Singh, Karan, et al.
Published: (2026)
by: Singh, Karan, et al.
Published: (2026)
Theoretical Foundation of Flow-Based Time Series Generation: Provable Approximation, Generalization, and Efficiency
by: Long, Jiangxuan, et al.
Published: (2025)
by: Long, Jiangxuan, et al.
Published: (2025)
Skywork-Math: Data Scaling Laws for Mathematical Reasoning in Large Language Models -- The Story Goes On
by: Zeng, Liang, et al.
Published: (2024)
by: Zeng, Liang, et al.
Published: (2024)
LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws
by: Mayilvahanan, Prasanna, et al.
Published: (2025)
by: Mayilvahanan, Prasanna, et al.
Published: (2025)
Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
by: Rafailov, Rafael, et al.
Published: (2024)
by: Rafailov, Rafael, et al.
Published: (2024)
Nemotron-Cascade: Scaling Cascaded Reinforcement Learning for General-Purpose Reasoning Models
by: Wang, Boxin, et al.
Published: (2025)
by: Wang, Boxin, et al.
Published: (2025)
Uncovering Scaling Laws for Large Language Models via Inverse Problems
by: Verma, Arun, et al.
Published: (2025)
by: Verma, Arun, et al.
Published: (2025)
Provable Scaling Laws for the Test-Time Compute of Large Language Models
by: Chen, Yanxi, et al.
Published: (2024)
by: Chen, Yanxi, et al.
Published: (2024)
From Self-Attention to Markov Models: Unveiling the Dynamics of Generative Transformers
by: Ildiz, M. Emrullah, et al.
Published: (2024)
by: Ildiz, M. Emrullah, et al.
Published: (2024)
(Mis)Fitting: A Survey of Scaling Laws
by: Li, Margaret, et al.
Published: (2025)
by: Li, Margaret, et al.
Published: (2025)
Finetune Once: Decoupling General & Domain Learning with Dynamic Boosted Annealing
by: Tang, Yang, et al.
Published: (2025)
by: Tang, Yang, et al.
Published: (2025)
Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient
by: Ludziejewski, Jan, et al.
Published: (2025)
by: Ludziejewski, Jan, et al.
Published: (2025)
Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws
by: Allen-Zhu, Zeyuan, et al.
Published: (2024)
by: Allen-Zhu, Zeyuan, et al.
Published: (2024)
Selecting Large Language Model to Fine-tune via Rectified Scaling Law
by: Lin, Haowei, et al.
Published: (2024)
by: Lin, Haowei, et al.
Published: (2024)
Enhancing Stochastic Gradient Descent: A Unified Framework and Novel Acceleration Methods for Faster Convergence
by: Deng, Yichuan, et al.
Published: (2024)
by: Deng, Yichuan, et al.
Published: (2024)
Similar Items
-
Unlocking the Theory Behind Scaling 1-Bit Neural Networks
by: Daliri, Majid, et al.
Published: (2024) -
Towards Infinite-Long Prefix in Transformer
by: Liang, Yingyu, et al.
Published: (2024) -
How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse
by: Deng, Yichuan, et al.
Published: (2024) -
Diversity of Transformer Layers: One Aspect of Parameter Scaling Laws
by: Kamigaito, Hidetaka, et al.
Published: (2025) -
Scaling Law with Learning Rate Annealing
by: Tissue, Howe, et al.
Published: (2024)