What Scales in Cross-Entropy Scaling Law?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yan, Junxi, Wei, Zixi, Ai, Qingyao, Liu, Yiqun, Zhan, Jingtao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Relative-Based Scaling Law for Neural Language Models
von: Yue, Baoqing, et al.
Veröffentlicht: (2025)
von: Yue, Baoqing, et al.
Veröffentlicht: (2025)
Option-ID Based Elimination For Multiple Choice Questions
von: Zhu, Zhenhao, et al.
Veröffentlicht: (2025)
von: Zhu, Zhenhao, et al.
Veröffentlicht: (2025)
Beyond Scalar Reward Model: Learning Generative Judge from Preference Data
von: Ye, Ziyi, et al.
Veröffentlicht: (2024)
von: Ye, Ziyi, et al.
Veröffentlicht: (2024)
Scaling Laws For Dense Retrieval
von: Fang, Yan, et al.
Veröffentlicht: (2024)
von: Fang, Yan, et al.
Veröffentlicht: (2024)
Capability-aware Prompt Reformulation Learning for Text-to-Image Generation
von: Zhan, Jingtao, et al.
Veröffentlicht: (2024)
von: Zhan, Jingtao, et al.
Veröffentlicht: (2024)
Distillation Scaling Laws
von: Busbridge, Dan, et al.
Veröffentlicht: (2025)
von: Busbridge, Dan, et al.
Veröffentlicht: (2025)
Evaluating Intelligence via Trial and Error
von: Zhan, Jingtao, et al.
Veröffentlicht: (2025)
von: Zhan, Jingtao, et al.
Veröffentlicht: (2025)
Exploring Scaling Laws for EHR Foundation Models
von: Zhang, Sheng, et al.
Veröffentlicht: (2025)
von: Zhang, Sheng, et al.
Veröffentlicht: (2025)
Scaling Law with Learning Rate Annealing
von: Tissue, Howe, et al.
Veröffentlicht: (2024)
von: Tissue, Howe, et al.
Veröffentlicht: (2024)
Skywork-Math: Data Scaling Laws for Mathematical Reasoning in Large Language Models -- The Story Goes On
von: Zeng, Liang, et al.
Veröffentlicht: (2024)
von: Zeng, Liang, et al.
Veröffentlicht: (2024)
To Memorize or to Retrieve: Scaling Laws for RAG-Considerate Pretraining
von: Singh, Karan, et al.
Veröffentlicht: (2026)
von: Singh, Karan, et al.
Veröffentlicht: (2026)
Can Language Models Discover Scaling Laws?
von: Lin, Haowei, et al.
Veröffentlicht: (2025)
von: Lin, Haowei, et al.
Veröffentlicht: (2025)
Theoretical Foundations of Scaling Law in Familial Models
von: Song, Huan, et al.
Veröffentlicht: (2025)
von: Song, Huan, et al.
Veröffentlicht: (2025)
Scaling Laws for Predicting Downstream Performance in LLMs
von: Chen, Yangyi, et al.
Veröffentlicht: (2024)
von: Chen, Yangyi, et al.
Veröffentlicht: (2024)
Scaling Laws for Fine-Grained Mixture of Experts
von: Krajewski, Jakub, et al.
Veröffentlicht: (2024)
von: Krajewski, Jakub, et al.
Veröffentlicht: (2024)
A Hitchhiker's Guide to Scaling Law Estimation
von: Choshen, Leshem, et al.
Veröffentlicht: (2024)
von: Choshen, Leshem, et al.
Veröffentlicht: (2024)
Batched Contextual Reinforcement: A Task-Scaling Law for Efficient Reasoning
von: Yang, Bangji, et al.
Veröffentlicht: (2026)
von: Yang, Bangji, et al.
Veröffentlicht: (2026)
What Matters for Model Merging at Scale?
von: Yadav, Prateek, et al.
Veröffentlicht: (2024)
von: Yadav, Prateek, et al.
Veröffentlicht: (2024)
P$^2$ Law: Scaling Law for Post-Training After Model Pruning
von: Chen, Xiaodong, et al.
Veröffentlicht: (2024)
von: Chen, Xiaodong, et al.
Veröffentlicht: (2024)
Reasoning: From Reflection to Solution
von: Li, Zixi
Veröffentlicht: (2025)
von: Li, Zixi
Veröffentlicht: (2025)
Entropy Centroids as Intrinsic Rewards for Test-Time Scaling
von: Zhao, Wenshuo, et al.
Veröffentlicht: (2026)
von: Zhao, Wenshuo, et al.
Veröffentlicht: (2026)
Towards a Comprehensive Scaling Law of Mixture-of-Experts
von: Zhao, Guoliang, et al.
Veröffentlicht: (2025)
von: Zhao, Guoliang, et al.
Veröffentlicht: (2025)
Predicting Task Performance with Context-aware Scaling Laws
von: Montgomery, Kyle, et al.
Veröffentlicht: (2025)
von: Montgomery, Kyle, et al.
Veröffentlicht: (2025)
Unifying Learning Dynamics and Generalization in Transformers Scaling Law
von: Yang, Chiwun
Veröffentlicht: (2025)
von: Yang, Chiwun
Veröffentlicht: (2025)
Loss-to-Loss Prediction: Scaling Laws for All Datasets
von: Brandfonbrener, David, et al.
Veröffentlicht: (2024)
von: Brandfonbrener, David, et al.
Veröffentlicht: (2024)
Observational Scaling Laws and the Predictability of Language Model Performance
von: Ruan, Yangjun, et al.
Veröffentlicht: (2024)
von: Ruan, Yangjun, et al.
Veröffentlicht: (2024)
LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws
von: Mayilvahanan, Prasanna, et al.
Veröffentlicht: (2025)
von: Mayilvahanan, Prasanna, et al.
Veröffentlicht: (2025)
Diversity of Transformer Layers: One Aspect of Parameter Scaling Laws
von: Kamigaito, Hidetaka, et al.
Veröffentlicht: (2025)
von: Kamigaito, Hidetaka, et al.
Veröffentlicht: (2025)
Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
von: Rafailov, Rafael, et al.
Veröffentlicht: (2024)
von: Rafailov, Rafael, et al.
Veröffentlicht: (2024)
(Mis)Fitting: A Survey of Scaling Laws
von: Li, Margaret, et al.
Veröffentlicht: (2025)
von: Li, Margaret, et al.
Veröffentlicht: (2025)
Query Augmentation by Decoding Semantics from Brain Signals
von: Ye, Ziyi, et al.
Veröffentlicht: (2024)
von: Ye, Ziyi, et al.
Veröffentlicht: (2024)
Task-Stratified Knowledge Scaling Laws for Post-Training Quantized Large Language Models
von: Zhou, Chenxi, et al.
Veröffentlicht: (2025)
von: Zhou, Chenxi, et al.
Veröffentlicht: (2025)
Foundations of GenIR
von: Ai, Qingyao, et al.
Veröffentlicht: (2025)
von: Ai, Qingyao, et al.
Veröffentlicht: (2025)
Uncovering Scaling Laws for Large Language Models via Inverse Problems
von: Verma, Arun, et al.
Veröffentlicht: (2025)
von: Verma, Arun, et al.
Veröffentlicht: (2025)
Provable Scaling Laws for the Test-Time Compute of Large Language Models
von: Chen, Yanxi, et al.
Veröffentlicht: (2024)
von: Chen, Yanxi, et al.
Veröffentlicht: (2024)
A Tale of Tails: Model Collapse as a Change of Scaling Laws
von: Dohmatob, Elvis, et al.
Veröffentlicht: (2024)
von: Dohmatob, Elvis, et al.
Veröffentlicht: (2024)
MathScale: Scaling Instruction Tuning for Mathematical Reasoning
von: Tang, Zhengyang, et al.
Veröffentlicht: (2024)
von: Tang, Zhengyang, et al.
Veröffentlicht: (2024)
OptScale: Probabilistic Optimality for Inference-time Scaling
von: Wang, Youkang, et al.
Veröffentlicht: (2025)
von: Wang, Youkang, et al.
Veröffentlicht: (2025)
Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient
von: Ludziejewski, Jan, et al.
Veröffentlicht: (2025)
von: Ludziejewski, Jan, et al.
Veröffentlicht: (2025)
Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws
von: Allen-Zhu, Zeyuan, et al.
Veröffentlicht: (2024)
von: Allen-Zhu, Zeyuan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Relative-Based Scaling Law for Neural Language Models
von: Yue, Baoqing, et al.
Veröffentlicht: (2025) -
Option-ID Based Elimination For Multiple Choice Questions
von: Zhu, Zhenhao, et al.
Veröffentlicht: (2025) -
Beyond Scalar Reward Model: Learning Generative Judge from Preference Data
von: Ye, Ziyi, et al.
Veröffentlicht: (2024) -
Scaling Laws For Dense Retrieval
von: Fang, Yan, et al.
Veröffentlicht: (2024) -
Capability-aware Prompt Reformulation Learning for Text-to-Image Generation
von: Zhan, Jingtao, et al.
Veröffentlicht: (2024)