Alternating Gradient Flows: A Theory of Feature Learning in Two-layer Neural Networks
Fuente:
arXiv
Guardado en:
| Autores principales: | Kunin, Daniel, Marchetti, Giovanni Luca, Chen, Feng, Karkada, Dhruva, Simon, James B., DeWeese, Michael R., Ganguli, Surya, Miolane, Nina |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Closed-Form Training Dynamics Reveal Learned Features and Linear Structure in Word2Vec-like Models
por: Karkada, Dhruva, et al.
Publicado: (2025)
por: Karkada, Dhruva, et al.
Publicado: (2025)
Revisiting Policy Gradients for Restricted Policy Classes: Escaping Myopic Local Optima with $k$-step Policy Gradients
por: DeWeese, Alex, et al.
Publicado: (2026)
por: DeWeese, Alex, et al.
Publicado: (2026)
A Theory of Saddle Escape in Deep Nonlinear Networks
por: Rawal, Divit, et al.
Publicado: (2026)
por: Rawal, Divit, et al.
Publicado: (2026)
Sequential Group Composition: A Window into the Mechanics of Deep Learning
por: Marchetti, Giovanni Luca, et al.
Publicado: (2026)
por: Marchetti, Giovanni Luca, et al.
Publicado: (2026)
The lazy (NTK) and rich ($μ$P) regimes: a gentle tutorial
por: Karkada, Dhruva
Publicado: (2024)
por: Karkada, Dhruva
Publicado: (2024)
Status Concerns and Library Professionalism
por: DeWeese, L. Carroll
Publicado: (1972)
por: DeWeese, L. Carroll
Publicado: (1972)
Stochastic Collapse: How Gradient Noise Attracts SGD Dynamics Towards Simpler Subnetworks
por: Chen, Feng, et al.
Publicado: (2023)
por: Chen, Feng, et al.
Publicado: (2023)
Thinking Beyond Visibility: A Near-Optimal Policy Framework for Locally Interdependent Multi-Agent MDPs
por: DeWeese, Alex, et al.
Publicado: (2025)
por: DeWeese, Alex, et al.
Publicado: (2025)
Locally Interdependent Multi-Agent MDP: Theoretical Framework for Decentralized Agents with Dynamic Dependencies
por: DeWeese, Alex, et al.
Publicado: (2024)
por: DeWeese, Alex, et al.
Publicado: (2024)
A Paradigm of Commitment: Toward Professional Identity for Librarians.
por: DeWeese, Lemuel Carroll, III
Publicado: (1970)
por: DeWeese, Lemuel Carroll, III
Publicado: (1970)
A Paradigm of Commitment
por: DeWeese, Lemuel Carroll, III
Publicado: (1970)
por: DeWeese, Lemuel Carroll, III
Publicado: (1970)
Beyond Linear Response: Equivalence between Thermodynamic Geometry and Optimal Transport
por: Zhong, Adrianne, et al.
Publicado: (2024)
por: Zhong, Adrianne, et al.
Publicado: (2024)
Temperature and flow data from a sediment tank experiment and numerical Advection-Dispersion Model code
por: Luce, Charles, et al.
Publicado: (2017)
por: Luce, Charles, et al.
Publicado: (2017)
Predicting kernel regression learning curves from only raw data statistics
por: Karkada, Dhruva, et al.
Publicado: (2025)
por: Karkada, Dhruva, et al.
Publicado: (2025)
More is Better in Modern Machine Learning: when Infinite Overparameterization is Optimal and Overfitting is Obligatory
por: Simon, James B., et al.
Publicado: (2023)
por: Simon, James B., et al.
Publicado: (2023)
Higher-order response theory in optimal stochastic thermodynamics
por: DAmbrosia, Samuel. H., et al.
Publicado: (2025)
por: DAmbrosia, Samuel. H., et al.
Publicado: (2025)
Time-Asymmetric Fluctuation Theorem and Efficient Free Energy Estimation
por: Zhong, Adrianne, et al.
Publicado: (2023)
por: Zhong, Adrianne, et al.
Publicado: (2023)
A General Framework for Robust G-Invariance in G-Equivariant Networks
por: Sanborn, Sophia, et al.
Publicado: (2023)
por: Sanborn, Sophia, et al.
Publicado: (2023)
On the Emergence of Linear Analogies in Word Embeddings
por: Korchinski, Daniel J., et al.
Publicado: (2025)
por: Korchinski, Daniel J., et al.
Publicado: (2025)
Get rich quick: exact solutions reveal how unbalanced initializations promote rapid feature learning
por: Kunin, Daniel, et al.
Publicado: (2024)
por: Kunin, Daniel, et al.
Publicado: (2024)
The Entropy of Floating-Point Numbers
por: Daniels, Sultan, et al.
Publicado: (2026)
por: Daniels, Sultan, et al.
Publicado: (2026)
There Will Be a Scientific Theory of Deep Learning
por: Simon, Jamie, et al.
Publicado: (2026)
por: Simon, Jamie, et al.
Publicado: (2026)
TopoTune : A Framework for Generalized Combinatorial Complex Neural Networks
por: Papillon, Mathilde, et al.
Publicado: (2024)
por: Papillon, Mathilde, et al.
Publicado: (2024)
The Thermodynamic Costs of Simple Linear Regression
por: D'Ambrosia, Samuel H., et al.
Publicado: (2026)
por: D'Ambrosia, Samuel H., et al.
Publicado: (2026)
An efficient algorithm for the Riemannian logarithm on the Stiefel manifold for a family of Riemannian metrics
por: Mataigne, Simon, et al.
Publicado: (2024)
por: Mataigne, Simon, et al.
Publicado: (2024)
Architectures of Topological Deep Learning: A Survey of Message-Passing Topological Neural Networks
por: Papillon, Mathilde, et al.
Publicado: (2023)
por: Papillon, Mathilde, et al.
Publicado: (2023)
Learning on a Razor's Edge: Identifiability and Singularity of Polynomial Neural Networks
por: Shahverdi, Vahid, et al.
Publicado: (2025)
por: Shahverdi, Vahid, et al.
Publicado: (2025)
Features are fate: a theory of transfer learning in high-dimensional regression
por: Tahir, Javan, et al.
Publicado: (2024)
por: Tahir, Javan, et al.
Publicado: (2024)
Symmetry in language statistics shapes the geometry of model representations
por: Karkada, Dhruva, et al.
Publicado: (2026)
por: Karkada, Dhruva, et al.
Publicado: (2026)
An analytic theory of creativity in convolutional diffusion models
por: Kamb, Mason, et al.
Publicado: (2024)
por: Kamb, Mason, et al.
Publicado: (2024)
Fooling LLM graders into giving better grades through neural activity guided adversarial prompting
por: Yamamura, Atsushi, et al.
Publicado: (2024)
por: Yamamura, Atsushi, et al.
Publicado: (2024)
Harmonics of Learning: Universal Fourier Features Emerge in Invariant Networks
por: Marchetti, Giovanni Luca, et al.
Publicado: (2023)
por: Marchetti, Giovanni Luca, et al.
Publicado: (2023)
The Selective Disk Bispectrum and Its Inversion, with Application to Multi-Reference Alignment
por: Myers, Adele, et al.
Publicado: (2025)
por: Myers, Adele, et al.
Publicado: (2025)
Meta-Learning for Better Learning: Using Meta-Learning Methods to Automatically Label Exam Questions with Detailed Learning Objectives
por: Zur, Amir, et al.
Publicado: (2023)
por: Zur, Amir, et al.
Publicado: (2023)
Rubik's Abstract Polytopes
por: Marchetti, Giovanni Luca
Publicado: (2025)
por: Marchetti, Giovanni Luca
Publicado: (2025)
Stochastic Gradient Descent for Two-layer Neural Networks
por: Cao, Dinghao, et al.
Publicado: (2024)
por: Cao, Dinghao, et al.
Publicado: (2024)
On the approximation of the Riemannian barycenter
por: Mataigne, Simon, et al.
Publicado: (2025)
por: Mataigne, Simon, et al.
Publicado: (2025)
Bounds on the geodesic distances on the Stiefel manifold for a family of Riemannian metrics
por: Mataigne, Simon, et al.
Publicado: (2024)
por: Mataigne, Simon, et al.
Publicado: (2024)
The Selective G-Bispectrum and its Inversion: Applications to G-Invariant Networks
por: Mataigne, Simon, et al.
Publicado: (2024)
por: Mataigne, Simon, et al.
Publicado: (2024)
Deriving Neural Scaling Laws from the statistics of natural language
por: Cagnetta, Francesco, et al.
Publicado: (2026)
por: Cagnetta, Francesco, et al.
Publicado: (2026)
Ejemplares similares
-
Closed-Form Training Dynamics Reveal Learned Features and Linear Structure in Word2Vec-like Models
por: Karkada, Dhruva, et al.
Publicado: (2025) -
Revisiting Policy Gradients for Restricted Policy Classes: Escaping Myopic Local Optima with $k$-step Policy Gradients
por: DeWeese, Alex, et al.
Publicado: (2026) -
A Theory of Saddle Escape in Deep Nonlinear Networks
por: Rawal, Divit, et al.
Publicado: (2026) -
Sequential Group Composition: A Window into the Mechanics of Deep Learning
por: Marchetti, Giovanni Luca, et al.
Publicado: (2026) -
The lazy (NTK) and rich ($μ$P) regimes: a gentle tutorial
por: Karkada, Dhruva
Publicado: (2024)