Saved in:
| Main Authors: | Simon, Jamie, Kunin, Daniel, Atanasov, Alexander, Boix-Adserà, Enric, Bordelon, Blake, Cohen, Jeremy, Ghosh, Nikhil, Guth, Florentin, Jacot, Arthur, Kamb, Mason, Karkada, Dhruva, Michaud, Eric J., Ottlik, Berkan, Turnbull, Joseph |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2604.21691 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Secret mixtures of experts inside your LLM
by: Boix-Adsera, Enric
Published: (2025)
by: Boix-Adsera, Enric
Published: (2025)
Towards a theory of model distillation
by: Boix-Adsera, Enric
Published: (2024)
by: Boix-Adsera, Enric
Published: (2024)
On the inductive bias of infinite-depth ResNets and the bottleneck rank
by: Boix-Adsera, Enric
Published: (2025)
by: Boix-Adsera, Enric
Published: (2025)
The power of fine-grained experts: Granularity boosts expressivity in Mixture of Experts
by: Boix-Adsera, Enric, et al.
Published: (2025)
by: Boix-Adsera, Enric, et al.
Published: (2025)
The lazy (NTK) and rich ($μ$P) regimes: a gentle tutorial
by: Karkada, Dhruva
Published: (2024)
by: Karkada, Dhruva
Published: (2024)
More is Better in Modern Machine Learning: when Infinite Overparameterization is Optimal and Overfitting is Obligatory
by: Simon, James B., et al.
Published: (2023)
by: Simon, James B., et al.
Published: (2023)
The merged-staircase property: a necessary and nearly sufficient condition for SGD learning of sparse functions on two-layer neural networks
by: Abbe, Emmanuel, et al.
Published: (2022)
by: Abbe, Emmanuel, et al.
Published: (2022)
Predicting kernel regression learning curves from only raw data statistics
by: Karkada, Dhruva, et al.
Published: (2025)
by: Karkada, Dhruva, et al.
Published: (2025)
How Feature Learning Can Improve Neural Scaling Laws
by: Bordelon, Blake, et al.
Published: (2024)
by: Bordelon, Blake, et al.
Published: (2024)
A Dynamical Model of Neural Scaling Laws
by: Bordelon, Blake, et al.
Published: (2024)
by: Bordelon, Blake, et al.
Published: (2024)
Toward universal steering and monitoring of AI models
by: Beaglehole, Daniel, et al.
Published: (2025)
by: Beaglehole, Daniel, et al.
Published: (2025)
The Features at Convergence Theorem: a first-principles alternative to the Neural Feature Ansatz for how networks learn representations
by: Boix-Adsera, Enric, et al.
Published: (2025)
by: Boix-Adsera, Enric, et al.
Published: (2025)
An analytic theory of creativity in convolutional diffusion models
by: Kamb, Mason, et al.
Published: (2024)
by: Kamb, Mason, et al.
Published: (2024)
Let Me Think! A Long Chain-of-Thought Can Be Worth Exponentially Many Short Ones
by: Mirtaheri, Parsa, et al.
Published: (2025)
by: Mirtaheri, Parsa, et al.
Published: (2025)
When can transformers reason with abstract symbols?
by: Boix-Adsera, Enric, et al.
Published: (2023)
by: Boix-Adsera, Enric, et al.
Published: (2023)
On the universality of neural encodings in CNNs
by: Guth, Florentin, et al.
Published: (2024)
by: Guth, Florentin, et al.
Published: (2024)
Classification-Denoising Networks
by: Thiry, Louis, et al.
Published: (2024)
by: Thiry, Louis, et al.
Published: (2024)
On the Emergence of Linear Analogies in Word Embeddings
by: Korchinski, Daniel J., et al.
Published: (2025)
by: Korchinski, Daniel J., et al.
Published: (2025)
Two-Point Deterministic Equivalence for Stochastic Gradient Dynamics in Linear Models
by: Atanasov, Alexander, et al.
Published: (2025)
by: Atanasov, Alexander, et al.
Published: (2025)
Alternating Gradient Flows: A Theory of Feature Learning in Two-layer Neural Networks
by: Kunin, Daniel, et al.
Published: (2025)
by: Kunin, Daniel, et al.
Published: (2025)
Theory of Optimal Learning Rate Schedules and Scaling Laws for a Random Feature Model
by: Bordelon, Blake, et al.
Published: (2026)
by: Bordelon, Blake, et al.
Published: (2026)
Disordered Dynamics in High Dimensions: Connections to Random Matrices and Machine Learning
by: Bordelon, Blake, et al.
Published: (2026)
by: Bordelon, Blake, et al.
Published: (2026)
Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer
by: Bordelon, Blake, et al.
Published: (2025)
by: Bordelon, Blake, et al.
Published: (2025)
Prompts have evil twins
by: Melamed, Rimon, et al.
Published: (2023)
by: Melamed, Rimon, et al.
Published: (2023)
Bottleneck Structure in Learned Features: Low-Dimension vs Regularity Tradeoff
by: Jacot, Arthur
Published: (2023)
by: Jacot, Arthur
Published: (2023)
Polynomial Speedup in Diffusion Models with the Multilevel Euler-Maruyama Method
by: Jacot, Arthur
Published: (2026)
by: Jacot, Arthur
Published: (2026)
Deep Learning as a Convex Paradigm of Computation: Minimizing Circuit Size with ResNets
by: Jacot, Arthur
Published: (2025)
by: Jacot, Arthur
Published: (2025)
Enhancing Motivation and Listening Comprehension Among Japanese EFL Students Through Translanguaging Practices
by: Blake Turnbull
Published: (2025)
by: Blake Turnbull
Published: (2025)
Annuaire de la société des nations, 1920 - 1927
by: Ottlik, Georges
Published: (1927)
by: Ottlik, Georges
Published: (1927)
Closed-Form Training Dynamics Reveal Learned Features and Linear Structure in Word2Vec-like Models
by: Karkada, Dhruva, et al.
Published: (2025)
by: Karkada, Dhruva, et al.
Published: (2025)
Symmetry in language statistics shapes the geometry of model representations
by: Karkada, Dhruva, et al.
Published: (2026)
by: Karkada, Dhruva, et al.
Published: (2026)
Learning normalized image densities via dual score matching
by: Guth, Florentin, et al.
Published: (2025)
by: Guth, Florentin, et al.
Published: (2025)
Xenillus Clypeator Robineau-Desvoidy and its Identity
by: Jacot, Arthur Paul
Published: (1929)
by: Jacot, Arthur Paul
Published: (1929)
Scientific publishing and social responsibility
by: Bernardo Turnbull
Published: (2019)
by: Bernardo Turnbull
Published: (2019)
Spectral Dynamics in Deep Networks: Feature Learning, Outlier Escape, and Learning Rate Transfer
by: Lauditi, Clarissa, et al.
Published: (2026)
by: Lauditi, Clarissa, et al.
Published: (2026)
Adaptive kernel predictors from feature-learning infinite limits of neural networks
by: Lauditi, Clarissa, et al.
Published: (2025)
by: Lauditi, Clarissa, et al.
Published: (2025)
Transfer Learning in Infinite Width Feature Learning Networks
by: Lauditi, Clarissa, et al.
Published: (2025)
by: Lauditi, Clarissa, et al.
Published: (2025)
The graphics of communication : typography, layout, design, production /^ Arthur T. Turnbull, Russell N. Baird
by: Turnbull, Arthur T
Published: (1980)
by: Turnbull, Arthur T
Published: (1980)
Learning Normalized Energy Models for Linear Inverse Problems
by: Zilberstein, Nicolas, et al.
Published: (2026)
by: Zilberstein, Nicolas, et al.
Published: (2026)
A Rainbow in Deep Network Black Boxes
by: Guth, Florentin, et al.
Published: (2023)
by: Guth, Florentin, et al.
Published: (2023)
Similar Items
-
Secret mixtures of experts inside your LLM
by: Boix-Adsera, Enric
Published: (2025) -
Towards a theory of model distillation
by: Boix-Adsera, Enric
Published: (2024) -
On the inductive bias of infinite-depth ResNets and the bottleneck rank
by: Boix-Adsera, Enric
Published: (2025) -
The power of fine-grained experts: Granularity boosts expressivity in Mixture of Experts
by: Boix-Adsera, Enric, et al.
Published: (2025) -
The lazy (NTK) and rich ($μ$P) regimes: a gentle tutorial
by: Karkada, Dhruva
Published: (2024)