Explaining Neural Scaling Laws
Fuente:
arXiv
Guardado en:
| Autores principales: | Bahri, Yasaman, Dyer, Ethan, Kaplan, Jared, Lee, Jaehoon, Sharma, Utkarsh |
|---|---|
| Formato: | Preprint |
| Publicado: |
2021
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Neural Scaling Laws Rooted in the Data Distribution
por: Brill, Ari
Publicado: (2024)
por: Brill, Ari
Publicado: (2024)
A Dynamical Model of Neural Scaling Laws
por: Bordelon, Blake, et al.
Publicado: (2024)
por: Bordelon, Blake, et al.
Publicado: (2024)
On the Emergence of Linear Analogies in Word Embeddings
por: Korchinski, Daniel J., et al.
Publicado: (2025)
por: Korchinski, Daniel J., et al.
Publicado: (2025)
How Feature Learning Can Improve Neural Scaling Laws
por: Bordelon, Blake, et al.
Publicado: (2024)
por: Bordelon, Blake, et al.
Publicado: (2024)
Symmetry in language statistics shapes the geometry of model representations
por: Karkada, Dhruva, et al.
Publicado: (2026)
por: Karkada, Dhruva, et al.
Publicado: (2026)
Theory of Scaling Laws for In-Context Regression: Depth, Width, Context and Time
por: Bordelon, Blake, et al.
Publicado: (2025)
por: Bordelon, Blake, et al.
Publicado: (2025)
The Quantization Model of Neural Scaling
por: Michaud, Eric J., et al.
Publicado: (2023)
por: Michaud, Eric J., et al.
Publicado: (2023)
No Free Lunch From Random Feature Ensembles: Scaling Laws and Near-Optimality Conditions
por: Ruben, Benjamin S., et al.
Publicado: (2024)
por: Ruben, Benjamin S., et al.
Publicado: (2024)
Theory of Optimal Learning Rate Schedules and Scaling Laws for a Random Feature Model
por: Bordelon, Blake, et al.
Publicado: (2026)
por: Bordelon, Blake, et al.
Publicado: (2026)
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures
por: Cagnetta, Francesco, et al.
Publicado: (2025)
por: Cagnetta, Francesco, et al.
Publicado: (2025)
Dropout Universality: Scaling Laws and Optimal Scheduling at the Edge-of-Chaos
por: Sarmiento, Lucas Fernandez
Publicado: (2026)
por: Sarmiento, Lucas Fernandez
Publicado: (2026)
Scaling Laws and Spectra of Shallow Neural Networks in the Feature Learning Regime
por: Defilippis, Leonardo, et al.
Publicado: (2025)
por: Defilippis, Leonardo, et al.
Publicado: (2025)
Universal Scaling Laws of Absorbing Phase Transitions in Artificial Deep Neural Networks
por: Tamai, Keiichi, et al.
Publicado: (2023)
por: Tamai, Keiichi, et al.
Publicado: (2023)
Asymmetric Scaling Laws from Sparse Features
por: Sous, John, et al.
Publicado: (2026)
por: Sous, John, et al.
Publicado: (2026)
Scaling and renormalization in high-dimensional regression
por: Atanasov, Alexander, et al.
Publicado: (2024)
por: Atanasov, Alexander, et al.
Publicado: (2024)
Formation of Representations in Neural Networks
por: Ziyin, Liu, et al.
Publicado: (2024)
por: Ziyin, Liu, et al.
Publicado: (2024)
Graph Neural Networks Do Not Always Oversmooth
por: Epping, Bastian, et al.
Publicado: (2024)
por: Epping, Bastian, et al.
Publicado: (2024)
A Boundary-Layer Mechanism for One-Third Scaling in Online Softmax Classification
por: Kühn, Marcel, et al.
Publicado: (2026)
por: Kühn, Marcel, et al.
Publicado: (2026)
From Kernels to Features: A Multi-Scale Adaptive Theory of Feature Learning
por: Rubin, Noa, et al.
Publicado: (2025)
por: Rubin, Noa, et al.
Publicado: (2025)
Exact Fixed-Point Constraints in Neural-ODEs with Provable Universality
por: Pacifico, Feliciano Giuseppe, et al.
Publicado: (2026)
por: Pacifico, Feliciano Giuseppe, et al.
Publicado: (2026)
Demolition and Reinforcement of Memories in Spin-Glass-like Neural Networks
por: Ventura, Enrico
Publicado: (2024)
por: Ventura, Enrico
Publicado: (2024)
Growing Neural Networks: Dynamic Evolution through Gradient Descent
por: Radhakrishnan, Anil, et al.
Publicado: (2025)
por: Radhakrishnan, Anil, et al.
Publicado: (2025)
The Rules-and-Facts Model for Simultaneous Generalization and Memorization in Neural Networks
por: Farné, Gabriele, et al.
Publicado: (2026)
por: Farné, Gabriele, et al.
Publicado: (2026)
Dynamical Mean-Field Theory of Self-Attention Neural Networks
por: Poc-López, Ángel, et al.
Publicado: (2024)
por: Poc-López, Ángel, et al.
Publicado: (2024)
Explaining the Machine Learning Solution of the Ising Model
por: Alamino, Roberto C.
Publicado: (2024)
por: Alamino, Roberto C.
Publicado: (2024)
Benchmarking Graph Neural Networks in Solving Hard Constraint Satisfaction Problems
por: Skenderi, Geri, et al.
Publicado: (2026)
por: Skenderi, Geri, et al.
Publicado: (2026)
A Random-Matrix Criterion for Initializing Gated Recurrent Neural Networks
por: Fioratti, Tommaso, et al.
Publicado: (2026)
por: Fioratti, Tommaso, et al.
Publicado: (2026)
Controlled Langevin Dynamics for Sampling of Feedforward Neural Networks Trained with Minibatches
por: Zambon, Alessandro, et al.
Publicado: (2026)
por: Zambon, Alessandro, et al.
Publicado: (2026)
A Federated Many-to-One Hopfield model for associative Neural Networks
por: Alessandrelli, Andrea, et al.
Publicado: (2026)
por: Alessandrelli, Andrea, et al.
Publicado: (2026)
Scaling Laws Do Not Scale
por: Diaz, Fernando, et al.
Publicado: (2023)
por: Diaz, Fernando, et al.
Publicado: (2023)
Single-Head Attention in High Dimensions: A Theory of Generalization, Weights Spectra, and Scaling Laws
por: Boncoraglio, Fabrizio, et al.
Publicado: (2025)
por: Boncoraglio, Fabrizio, et al.
Publicado: (2025)
Kernel Renormalization in Bayesian Deep Neural Networks: the Equivalent Wishart Ansatz in the Proportional Regime
por: Baglioni, Paolo, et al.
Publicado: (2026)
por: Baglioni, Paolo, et al.
Publicado: (2026)
Graph Neural Network Approach to Predicting Magnetization in Quasi-One-Dimensional Ising Systems
por: Slavin, V., et al.
Publicado: (2025)
por: Slavin, V., et al.
Publicado: (2025)
Siamese Neural Network for Label-Efficient Critical Phenomena Prediction in 3D Percolation Models
por: Wang, Shanshan, et al.
Publicado: (2025)
por: Wang, Shanshan, et al.
Publicado: (2025)
Statistical Physics of Deep Neural Networks: Generalization Capability, Beyond the Infinite Width, and Feature Learning
por: Ariosto, Sebastiano
Publicado: (2025)
por: Ariosto, Sebastiano
Publicado: (2025)
Deep Learning as Neural Low-Degree Filtering: A Spectral Theory of Hierarchical Feature Learning
por: Dandi, Yatin, et al.
Publicado: (2026)
por: Dandi, Yatin, et al.
Publicado: (2026)
Explaining the effects of non-convergent sampling in the training of Energy-Based Models
por: Agoritsas, Elisabeth, et al.
Publicado: (2023)
por: Agoritsas, Elisabeth, et al.
Publicado: (2023)
The RL Perceptron: Generalisation Dynamics of Policy Learning in High Dimensions
por: Patel, Nishil, et al.
Publicado: (2023)
por: Patel, Nishil, et al.
Publicado: (2023)
An exactly solvable model for emergence and scaling laws in the multitask sparse parity problem
por: Nam, Yoonsoo, et al.
Publicado: (2024)
por: Nam, Yoonsoo, et al.
Publicado: (2024)
Dataset-Free Weight-Initialization on Restricted Boltzmann Machine
por: Yasuda, Muneki, et al.
Publicado: (2024)
por: Yasuda, Muneki, et al.
Publicado: (2024)
Ejemplares similares
-
Neural Scaling Laws Rooted in the Data Distribution
por: Brill, Ari
Publicado: (2024) -
A Dynamical Model of Neural Scaling Laws
por: Bordelon, Blake, et al.
Publicado: (2024) -
On the Emergence of Linear Analogies in Word Embeddings
por: Korchinski, Daniel J., et al.
Publicado: (2025) -
How Feature Learning Can Improve Neural Scaling Laws
por: Bordelon, Blake, et al.
Publicado: (2024) -
Symmetry in language statistics shapes the geometry of model representations
por: Karkada, Dhruva, et al.
Publicado: (2026)