Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws
Fuente:
arXiv
Salvato in:
| Autori principali: | Allen-Zhu, Zeyuan, Li, Yuanzhi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Physics of Language Models: Part 3.2, Knowledge Manipulation
di: Allen-Zhu, Zeyuan, et al.
Pubblicazione: (2023)
di: Allen-Zhu, Zeyuan, et al.
Pubblicazione: (2023)
Physics of Language Models: Part 3.1, Knowledge Storage and Extraction
di: Allen-Zhu, Zeyuan, et al.
Pubblicazione: (2023)
di: Allen-Zhu, Zeyuan, et al.
Pubblicazione: (2023)
Physics of Language Models: Part 1, Learning Hierarchical Language Structures
di: Allen-Zhu, Zeyuan, et al.
Pubblicazione: (2023)
di: Allen-Zhu, Zeyuan, et al.
Pubblicazione: (2023)
Physics of Language Models: Part 2.1, Grade-School Math and the Hidden Reasoning Process
di: Ye, Tian, et al.
Pubblicazione: (2024)
di: Ye, Tian, et al.
Pubblicazione: (2024)
Physics of Language Models: Part 2.2, How to Learn From Mistakes on Grade-School Math Problems
di: Ye, Tian, et al.
Pubblicazione: (2024)
di: Ye, Tian, et al.
Pubblicazione: (2024)
Towards the Law of Capacity Gap in Distilling Language Models
di: Zhang, Chen, et al.
Pubblicazione: (2023)
di: Zhang, Chen, et al.
Pubblicazione: (2023)
Task-Stratified Knowledge Scaling Laws for Post-Training Quantized Large Language Models
di: Zhou, Chenxi, et al.
Pubblicazione: (2025)
di: Zhou, Chenxi, et al.
Pubblicazione: (2025)
Can Language Models Discover Scaling Laws?
di: Lin, Haowei, et al.
Pubblicazione: (2025)
di: Lin, Haowei, et al.
Pubblicazione: (2025)
Observational Scaling Laws and the Predictability of Language Model Performance
di: Ruan, Yangjun, et al.
Pubblicazione: (2024)
di: Ruan, Yangjun, et al.
Pubblicazione: (2024)
Relative-Based Scaling Law for Neural Language Models
di: Yue, Baoqing, et al.
Pubblicazione: (2025)
di: Yue, Baoqing, et al.
Pubblicazione: (2025)
Provable Scaling Laws for the Test-Time Compute of Large Language Models
di: Chen, Yanxi, et al.
Pubblicazione: (2024)
di: Chen, Yanxi, et al.
Pubblicazione: (2024)
Selecting Large Language Model to Fine-tune via Rectified Scaling Law
di: Lin, Haowei, et al.
Pubblicazione: (2024)
di: Lin, Haowei, et al.
Pubblicazione: (2024)
Uncovering Scaling Laws for Large Language Models via Inverse Problems
di: Verma, Arun, et al.
Pubblicazione: (2025)
di: Verma, Arun, et al.
Pubblicazione: (2025)
Theoretical Foundations of Scaling Law in Familial Models
di: Song, Huan, et al.
Pubblicazione: (2025)
di: Song, Huan, et al.
Pubblicazione: (2025)
P$^2$ Law: Scaling Law for Post-Training After Model Pruning
di: Chen, Xiaodong, et al.
Pubblicazione: (2024)
di: Chen, Xiaodong, et al.
Pubblicazione: (2024)
SE-Bench: Benchmarking Self-Evolution with Knowledge Internalization
di: Yuan, Jiarui, et al.
Pubblicazione: (2026)
di: Yuan, Jiarui, et al.
Pubblicazione: (2026)
Exploring Scaling Laws for EHR Foundation Models
di: Zhang, Sheng, et al.
Pubblicazione: (2025)
di: Zhang, Sheng, et al.
Pubblicazione: (2025)
Skywork-Math: Data Scaling Laws for Mathematical Reasoning in Large Language Models -- The Story Goes On
di: Zeng, Liang, et al.
Pubblicazione: (2024)
di: Zeng, Liang, et al.
Pubblicazione: (2024)
Distillation Scaling Laws
di: Busbridge, Dan, et al.
Pubblicazione: (2025)
di: Busbridge, Dan, et al.
Pubblicazione: (2025)
Evaluating the Factuality of Large Language Models using Large-Scale Knowledge Graphs
di: Liu, Xiaoze, et al.
Pubblicazione: (2024)
di: Liu, Xiaoze, et al.
Pubblicazione: (2024)
What Scales in Cross-Entropy Scaling Law?
di: Yan, Junxi, et al.
Pubblicazione: (2025)
di: Yan, Junxi, et al.
Pubblicazione: (2025)
Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
di: Rafailov, Rafael, et al.
Pubblicazione: (2024)
di: Rafailov, Rafael, et al.
Pubblicazione: (2024)
Rethinking LLM-as-a-Judge: Representation-as-a-Judge with Small Language Models via Semantic Capacity Asymmetry
di: Li, Zhuochun, et al.
Pubblicazione: (2026)
di: Li, Zhuochun, et al.
Pubblicazione: (2026)
A Tale of Tails: Model Collapse as a Change of Scaling Laws
di: Dohmatob, Elvis, et al.
Pubblicazione: (2024)
di: Dohmatob, Elvis, et al.
Pubblicazione: (2024)
Scaling Law with Learning Rate Annealing
di: Tissue, Howe, et al.
Pubblicazione: (2024)
di: Tissue, Howe, et al.
Pubblicazione: (2024)
Scaling Embeddings Outperforms Scaling Experts in Language Models
di: Liu, Hong, et al.
Pubblicazione: (2026)
di: Liu, Hong, et al.
Pubblicazione: (2026)
BiMix: A Bivariate Data Mixing Law for Language Model Pretraining
di: Ge, Ce, et al.
Pubblicazione: (2024)
di: Ge, Ce, et al.
Pubblicazione: (2024)
Towards a Comprehensive Scaling Law of Mixture-of-Experts
di: Zhao, Guoliang, et al.
Pubblicazione: (2025)
di: Zhao, Guoliang, et al.
Pubblicazione: (2025)
Scaling over Scaling: Exploring Test-Time Scaling Plateau in Large Reasoning Models
di: Wang, Jian, et al.
Pubblicazione: (2025)
di: Wang, Jian, et al.
Pubblicazione: (2025)
Scaling Laws for Predicting Downstream Performance in LLMs
di: Chen, Yangyi, et al.
Pubblicazione: (2024)
di: Chen, Yangyi, et al.
Pubblicazione: (2024)
Scaling Laws for Fine-Grained Mixture of Experts
di: Krajewski, Jakub, et al.
Pubblicazione: (2024)
di: Krajewski, Jakub, et al.
Pubblicazione: (2024)
A Hitchhiker's Guide to Scaling Law Estimation
di: Choshen, Leshem, et al.
Pubblicazione: (2024)
di: Choshen, Leshem, et al.
Pubblicazione: (2024)
(Mis)Fitting: A Survey of Scaling Laws
di: Li, Margaret, et al.
Pubblicazione: (2025)
di: Li, Margaret, et al.
Pubblicazione: (2025)
A Law of Next-Token Prediction in Large Language Models
di: He, Hangfeng, et al.
Pubblicazione: (2024)
di: He, Hangfeng, et al.
Pubblicazione: (2024)
Knowledge Fusion of Large Language Models Via Modular SkillPacks
di: Du, Guodong, et al.
Pubblicazione: (2025)
di: Du, Guodong, et al.
Pubblicazione: (2025)
Training Language Models with Language Feedback at Scale
di: Scheurer, Jérémy, et al.
Pubblicazione: (2023)
di: Scheurer, Jérémy, et al.
Pubblicazione: (2023)
Loss-to-Loss Prediction: Scaling Laws for All Datasets
di: Brandfonbrener, David, et al.
Pubblicazione: (2024)
di: Brandfonbrener, David, et al.
Pubblicazione: (2024)
Predicting Task Performance with Context-aware Scaling Laws
di: Montgomery, Kyle, et al.
Pubblicazione: (2025)
di: Montgomery, Kyle, et al.
Pubblicazione: (2025)
Unifying Learning Dynamics and Generalization in Transformers Scaling Law
di: Yang, Chiwun
Pubblicazione: (2025)
di: Yang, Chiwun
Pubblicazione: (2025)
To Memorize or to Retrieve: Scaling Laws for RAG-Considerate Pretraining
di: Singh, Karan, et al.
Pubblicazione: (2026)
di: Singh, Karan, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Physics of Language Models: Part 3.2, Knowledge Manipulation
di: Allen-Zhu, Zeyuan, et al.
Pubblicazione: (2023) -
Physics of Language Models: Part 3.1, Knowledge Storage and Extraction
di: Allen-Zhu, Zeyuan, et al.
Pubblicazione: (2023) -
Physics of Language Models: Part 1, Learning Hierarchical Language Structures
di: Allen-Zhu, Zeyuan, et al.
Pubblicazione: (2023) -
Physics of Language Models: Part 2.1, Grade-School Math and the Hidden Reasoning Process
di: Ye, Tian, et al.
Pubblicazione: (2024) -
Physics of Language Models: Part 2.2, How to Learn From Mistakes on Grade-School Math Problems
di: Ye, Tian, et al.
Pubblicazione: (2024)