Deriving Neural Scaling Laws from the statistics of natural language
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cagnetta, Francesco, Raventós, Allan, Ganguli, Surya, Wyart, Matthieu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards a theory of how the structure of language is acquired by deep neural networks
von: Cagnetta, Francesco, et al.
Veröffentlicht: (2024)
von: Cagnetta, Francesco, et al.
Veröffentlicht: (2024)
Deep networks learn to parse uniform-depth context-free languages from local statistics
von: Parley, Jack T., et al.
Veröffentlicht: (2026)
von: Parley, Jack T., et al.
Veröffentlicht: (2026)
Rethinking Fine-Tuning when Scaling Test-Time Compute: Limiting Confidence Improves Mathematical Reasoning
von: Chen, Feng, et al.
Veröffentlicht: (2025)
von: Chen, Feng, et al.
Veröffentlicht: (2025)
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures
von: Cagnetta, Francesco, et al.
Veröffentlicht: (2025)
von: Cagnetta, Francesco, et al.
Veröffentlicht: (2025)
Get rich quick: exact solutions reveal how unbalanced initializations promote rapid feature learning
von: Kunin, Daniel, et al.
Veröffentlicht: (2024)
von: Kunin, Daniel, et al.
Veröffentlicht: (2024)
Learning curves theory for hierarchically compositional data with power-law distributed features
von: Cagnetta, Francesco, et al.
Veröffentlicht: (2025)
von: Cagnetta, Francesco, et al.
Veröffentlicht: (2025)
How Compositional Generalization and Creativity Improve as Diffusion Models are Trained
von: Favero, Alessandro, et al.
Veröffentlicht: (2025)
von: Favero, Alessandro, et al.
Veröffentlicht: (2025)
How Deep Neural Networks Learn Compositional Data: The Random Hierarchy Model
von: Cagnetta, Francesco, et al.
Veröffentlicht: (2023)
von: Cagnetta, Francesco, et al.
Veröffentlicht: (2023)
TFGN: Task-Free, Replay-Free Continual Pre-Training Without Catastrophic Forgetting at LLM Scale
von: Ganguli, Anurup
Veröffentlicht: (2026)
von: Ganguli, Anurup
Veröffentlicht: (2026)
Stochastic Collapse: How Gradient Noise Attracts SGD Dynamics Towards Simpler Subnetworks
von: Chen, Feng, et al.
Veröffentlicht: (2023)
von: Chen, Feng, et al.
Veröffentlicht: (2023)
From Kepler to Newton: Inductive Biases Guide Learned World Models in Transformers
von: Liu, Ziming, et al.
Veröffentlicht: (2026)
von: Liu, Ziming, et al.
Veröffentlicht: (2026)
An analytic theory of creativity in convolutional diffusion models
von: Kamb, Mason, et al.
Veröffentlicht: (2024)
von: Kamb, Mason, et al.
Veröffentlicht: (2024)
Contrastive Concept-Tree Search for LLM-Assisted Algorithm Discovery
von: Leleu, Timothee, et al.
Veröffentlicht: (2026)
von: Leleu, Timothee, et al.
Veröffentlicht: (2026)
On the Optimizer Dependence of Neural Scaling Laws
von: Ramani, Vansh, et al.
Veröffentlicht: (2026)
von: Ramani, Vansh, et al.
Veröffentlicht: (2026)
Towards Neural Scaling Laws on Graphs
von: Liu, Jingzhe, et al.
Veröffentlicht: (2024)
von: Liu, Jingzhe, et al.
Veröffentlicht: (2024)
Information-Theoretic Foundations for Neural Scaling Laws
von: Jeon, Hong Jun, et al.
Veröffentlicht: (2024)
von: Jeon, Hong Jun, et al.
Veröffentlicht: (2024)
A Hierarchical Language Model with Predictable Scaling Laws and Provable Benefits of Reasoning
von: Gaitonde, Jason, et al.
Veröffentlicht: (2026)
von: Gaitonde, Jason, et al.
Veröffentlicht: (2026)
Towards Neural Scaling Laws for Time Series Foundation Models
von: Yao, Qingren, et al.
Veröffentlicht: (2024)
von: Yao, Qingren, et al.
Veröffentlicht: (2024)
Bayesian Neural Scaling Law Extrapolation with Prior-Data Fitted Networks
von: Lee, Dongwoo, et al.
Veröffentlicht: (2025)
von: Lee, Dongwoo, et al.
Veröffentlicht: (2025)
Do Neural Scaling Laws Exist on Graph Self-Supervised Learning?
von: Ma, Qian, et al.
Veröffentlicht: (2024)
von: Ma, Qian, et al.
Veröffentlicht: (2024)
Scaling Laws and Symmetry, Evidence from Neural Force Fields
von: Ngo, Khang, et al.
Veröffentlicht: (2025)
von: Ngo, Khang, et al.
Veröffentlicht: (2025)
Derived Fields Preserve Fine-Scale Detail in Budgeted Neural Simulators
von: Wang, Wenshuo, et al.
Veröffentlicht: (2026)
von: Wang, Wenshuo, et al.
Veröffentlicht: (2026)
MARK: Memory Augmented Refinement of Knowledge
von: Ganguli, Anish, et al.
Veröffentlicht: (2025)
von: Ganguli, Anish, et al.
Veröffentlicht: (2025)
Effective Frontiers: A Unification of Neural Scaling Laws
von: Zou, Jiaxuan, et al.
Veröffentlicht: (2026)
von: Zou, Jiaxuan, et al.
Veröffentlicht: (2026)
Relative-Based Scaling Law for Neural Language Models
von: Yue, Baoqing, et al.
Veröffentlicht: (2025)
von: Yue, Baoqing, et al.
Veröffentlicht: (2025)
A Resource Model For Neural Scaling Law
von: Song, Jinyeop, et al.
Veröffentlicht: (2024)
von: Song, Jinyeop, et al.
Veröffentlicht: (2024)
The Neural Pruning Law Hypothesis
von: Barbulescu, Eugen, et al.
Veröffentlicht: (2025)
von: Barbulescu, Eugen, et al.
Veröffentlicht: (2025)
Efficiency optimization of large-scale language models based on deep learning in natural language processing tasks
von: Mei, Taiyuan, et al.
Veröffentlicht: (2024)
von: Mei, Taiyuan, et al.
Veröffentlicht: (2024)
Universal Neural Functionals
von: Zhou, Allan, et al.
Veröffentlicht: (2024)
von: Zhou, Allan, et al.
Veröffentlicht: (2024)
On Implications of Scaling Laws on Feature Superposition
von: Katta, Pavan
Veröffentlicht: (2024)
von: Katta, Pavan
Veröffentlicht: (2024)
Scaling Law Hypothesis for Multimodal Model
von: Sun, Qingyun, et al.
Veröffentlicht: (2024)
von: Sun, Qingyun, et al.
Veröffentlicht: (2024)
Scaling Law for Time Series Forecasting
von: Shi, Jingzhe, et al.
Veröffentlicht: (2024)
von: Shi, Jingzhe, et al.
Veröffentlicht: (2024)
Symmetry in language statistics shapes the geometry of model representations
von: Karkada, Dhruva, et al.
Veröffentlicht: (2026)
von: Karkada, Dhruva, et al.
Veröffentlicht: (2026)
Wukong: Towards a Scaling Law for Large-Scale Recommendation
von: Zhang, Buyun, et al.
Veröffentlicht: (2024)
von: Zhang, Buyun, et al.
Veröffentlicht: (2024)
Time Matters: Scaling Laws for Any Budget
von: Inbar, Itay, et al.
Veröffentlicht: (2024)
von: Inbar, Itay, et al.
Veröffentlicht: (2024)
Provable Scaling Laws of Feature Emergence from Learning Dynamics of Grokking
von: Tian, Yuandong
Veröffentlicht: (2025)
von: Tian, Yuandong
Veröffentlicht: (2025)
Purity Law for Generalizable Neural TSP Solvers
von: Liu, Wenzhao, et al.
Veröffentlicht: (2025)
von: Liu, Wenzhao, et al.
Veröffentlicht: (2025)
A Neural Scaling Law from Lottery Ticket Ensembling
von: Liu, Ziming, et al.
Veröffentlicht: (2023)
von: Liu, Ziming, et al.
Veröffentlicht: (2023)
Scaling Laws for Precision in High-Dimensional Linear Regression
von: Zhang, Dechen, et al.
Veröffentlicht: (2026)
von: Zhang, Dechen, et al.
Veröffentlicht: (2026)
Improving the Scaling Laws of Synthetic Data with Deliberate Practice
von: Askari-Hemmat, Reyhane, et al.
Veröffentlicht: (2025)
von: Askari-Hemmat, Reyhane, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Towards a theory of how the structure of language is acquired by deep neural networks
von: Cagnetta, Francesco, et al.
Veröffentlicht: (2024) -
Deep networks learn to parse uniform-depth context-free languages from local statistics
von: Parley, Jack T., et al.
Veröffentlicht: (2026) -
Rethinking Fine-Tuning when Scaling Test-Time Compute: Limiting Confidence Improves Mathematical Reasoning
von: Chen, Feng, et al.
Veröffentlicht: (2025) -
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures
von: Cagnetta, Francesco, et al.
Veröffentlicht: (2025) -
Get rich quick: exact solutions reveal how unbalanced initializations promote rapid feature learning
von: Kunin, Daniel, et al.
Veröffentlicht: (2024)