There Will Be a Scientific Theory of Deep Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Simon, Jamie, Kunin, Daniel, Atanasov, Alexander, Boix-Adserà, Enric, Bordelon, Blake, Cohen, Jeremy, Ghosh, Nikhil, Guth, Florentin, Jacot, Arthur, Kamb, Mason, Karkada, Dhruva, Michaud, Eric J., Ottlik, Berkan, Turnbull, Joseph |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Secret mixtures of experts inside your LLM
von: Boix-Adsera, Enric
Veröffentlicht: (2025)
von: Boix-Adsera, Enric
Veröffentlicht: (2025)
Towards a theory of model distillation
von: Boix-Adsera, Enric
Veröffentlicht: (2024)
von: Boix-Adsera, Enric
Veröffentlicht: (2024)
On the inductive bias of infinite-depth ResNets and the bottleneck rank
von: Boix-Adsera, Enric
Veröffentlicht: (2025)
von: Boix-Adsera, Enric
Veröffentlicht: (2025)
The power of fine-grained experts: Granularity boosts expressivity in Mixture of Experts
von: Boix-Adsera, Enric, et al.
Veröffentlicht: (2025)
von: Boix-Adsera, Enric, et al.
Veröffentlicht: (2025)
The lazy (NTK) and rich ($μ$P) regimes: a gentle tutorial
von: Karkada, Dhruva
Veröffentlicht: (2024)
von: Karkada, Dhruva
Veröffentlicht: (2024)
More is Better in Modern Machine Learning: when Infinite Overparameterization is Optimal and Overfitting is Obligatory
von: Simon, James B., et al.
Veröffentlicht: (2023)
von: Simon, James B., et al.
Veröffentlicht: (2023)
Predicting kernel regression learning curves from only raw data statistics
von: Karkada, Dhruva, et al.
Veröffentlicht: (2025)
von: Karkada, Dhruva, et al.
Veröffentlicht: (2025)
The merged-staircase property: a necessary and nearly sufficient condition for SGD learning of sparse functions on two-layer neural networks
von: Abbe, Emmanuel, et al.
Veröffentlicht: (2022)
von: Abbe, Emmanuel, et al.
Veröffentlicht: (2022)
How Feature Learning Can Improve Neural Scaling Laws
von: Bordelon, Blake, et al.
Veröffentlicht: (2024)
von: Bordelon, Blake, et al.
Veröffentlicht: (2024)
A Dynamical Model of Neural Scaling Laws
von: Bordelon, Blake, et al.
Veröffentlicht: (2024)
von: Bordelon, Blake, et al.
Veröffentlicht: (2024)
Toward universal steering and monitoring of AI models
von: Beaglehole, Daniel, et al.
Veröffentlicht: (2025)
von: Beaglehole, Daniel, et al.
Veröffentlicht: (2025)
An analytic theory of creativity in convolutional diffusion models
von: Kamb, Mason, et al.
Veröffentlicht: (2024)
von: Kamb, Mason, et al.
Veröffentlicht: (2024)
The Features at Convergence Theorem: a first-principles alternative to the Neural Feature Ansatz for how networks learn representations
von: Boix-Adsera, Enric, et al.
Veröffentlicht: (2025)
von: Boix-Adsera, Enric, et al.
Veröffentlicht: (2025)
Let Me Think! A Long Chain-of-Thought Can Be Worth Exponentially Many Short Ones
von: Mirtaheri, Parsa, et al.
Veröffentlicht: (2025)
von: Mirtaheri, Parsa, et al.
Veröffentlicht: (2025)
On the universality of neural encodings in CNNs
von: Guth, Florentin, et al.
Veröffentlicht: (2024)
von: Guth, Florentin, et al.
Veröffentlicht: (2024)
Classification-Denoising Networks
von: Thiry, Louis, et al.
Veröffentlicht: (2024)
von: Thiry, Louis, et al.
Veröffentlicht: (2024)
When can transformers reason with abstract symbols?
von: Boix-Adsera, Enric, et al.
Veröffentlicht: (2023)
von: Boix-Adsera, Enric, et al.
Veröffentlicht: (2023)
On the Emergence of Linear Analogies in Word Embeddings
von: Korchinski, Daniel J., et al.
Veröffentlicht: (2025)
von: Korchinski, Daniel J., et al.
Veröffentlicht: (2025)
Bottleneck Structure in Learned Features: Low-Dimension vs Regularity Tradeoff
von: Jacot, Arthur
Veröffentlicht: (2023)
von: Jacot, Arthur
Veröffentlicht: (2023)
Polynomial Speedup in Diffusion Models with the Multilevel Euler-Maruyama Method
von: Jacot, Arthur
Veröffentlicht: (2026)
von: Jacot, Arthur
Veröffentlicht: (2026)
Deep Learning as a Convex Paradigm of Computation: Minimizing Circuit Size with ResNets
von: Jacot, Arthur
Veröffentlicht: (2025)
von: Jacot, Arthur
Veröffentlicht: (2025)
Two-Point Deterministic Equivalence for Stochastic Gradient Dynamics in Linear Models
von: Atanasov, Alexander, et al.
Veröffentlicht: (2025)
von: Atanasov, Alexander, et al.
Veröffentlicht: (2025)
Theory of Optimal Learning Rate Schedules and Scaling Laws for a Random Feature Model
von: Bordelon, Blake, et al.
Veröffentlicht: (2026)
von: Bordelon, Blake, et al.
Veröffentlicht: (2026)
Disordered Dynamics in High Dimensions: Connections to Random Matrices and Machine Learning
von: Bordelon, Blake, et al.
Veröffentlicht: (2026)
von: Bordelon, Blake, et al.
Veröffentlicht: (2026)
Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer
von: Bordelon, Blake, et al.
Veröffentlicht: (2025)
von: Bordelon, Blake, et al.
Veröffentlicht: (2025)
Annuaire de la société des nations, 1920 - 1927
von: Ottlik, Georges
Veröffentlicht: (1927)
von: Ottlik, Georges
Veröffentlicht: (1927)
Enhancing Motivation and Listening Comprehension Among Japanese EFL Students Through Translanguaging Practices
von: Blake Turnbull
Veröffentlicht: (2025)
von: Blake Turnbull
Veröffentlicht: (2025)
Alternating Gradient Flows: A Theory of Feature Learning in Two-layer Neural Networks
von: Kunin, Daniel, et al.
Veröffentlicht: (2025)
von: Kunin, Daniel, et al.
Veröffentlicht: (2025)
Xenillus Clypeator Robineau-Desvoidy and its Identity
von: Jacot, Arthur Paul
Veröffentlicht: (1929)
von: Jacot, Arthur Paul
Veröffentlicht: (1929)
Prompts have evil twins
von: Melamed, Rimon, et al.
Veröffentlicht: (2023)
von: Melamed, Rimon, et al.
Veröffentlicht: (2023)
Closed-Form Training Dynamics Reveal Learned Features and Linear Structure in Word2Vec-like Models
von: Karkada, Dhruva, et al.
Veröffentlicht: (2025)
von: Karkada, Dhruva, et al.
Veröffentlicht: (2025)
Symmetry in language statistics shapes the geometry of model representations
von: Karkada, Dhruva, et al.
Veröffentlicht: (2026)
von: Karkada, Dhruva, et al.
Veröffentlicht: (2026)
Scientific publishing and social responsibility
von: Bernardo Turnbull
Veröffentlicht: (2019)
von: Bernardo Turnbull
Veröffentlicht: (2019)
Learning normalized image densities via dual score matching
von: Guth, Florentin, et al.
Veröffentlicht: (2025)
von: Guth, Florentin, et al.
Veröffentlicht: (2025)
The graphics of communication : typography, layout, design, production /^ Arthur T. Turnbull, Russell N. Baird
von: Turnbull, Arthur T
Veröffentlicht: (1980)
von: Turnbull, Arthur T
Veröffentlicht: (1980)
Spectral Dynamics in Deep Networks: Feature Learning, Outlier Escape, and Learning Rate Transfer
von: Lauditi, Clarissa, et al.
Veröffentlicht: (2026)
von: Lauditi, Clarissa, et al.
Veröffentlicht: (2026)
Adaptive kernel predictors from feature-learning infinite limits of neural networks
von: Lauditi, Clarissa, et al.
Veröffentlicht: (2025)
von: Lauditi, Clarissa, et al.
Veröffentlicht: (2025)
Transfer Learning in Infinite Width Feature Learning Networks
von: Lauditi, Clarissa, et al.
Veröffentlicht: (2025)
von: Lauditi, Clarissa, et al.
Veröffentlicht: (2025)
Hamiltonian Mechanics of Feature Learning: Bottleneck Structure in Leaky ResNets
von: Jacot, Arthur, et al.
Veröffentlicht: (2024)
von: Jacot, Arthur, et al.
Veröffentlicht: (2024)
Which Frequencies do CNNs Need? Emergent Bottleneck Structure in Feature Learning
von: Wen, Yuxiao, et al.
Veröffentlicht: (2024)
von: Wen, Yuxiao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Secret mixtures of experts inside your LLM
von: Boix-Adsera, Enric
Veröffentlicht: (2025) -
Towards a theory of model distillation
von: Boix-Adsera, Enric
Veröffentlicht: (2024) -
On the inductive bias of infinite-depth ResNets and the bottleneck rank
von: Boix-Adsera, Enric
Veröffentlicht: (2025) -
The power of fine-grained experts: Granularity boosts expressivity in Mixture of Experts
von: Boix-Adsera, Enric, et al.
Veröffentlicht: (2025) -
The lazy (NTK) and rich ($μ$P) regimes: a gentle tutorial
von: Karkada, Dhruva
Veröffentlicht: (2024)