How Deep Networks Learn Sparse and Hierarchical Data: the Sparse Random Hierarchy Model
Fuente:
arXiv
Saved in:
| Main Authors: | Tomasini, Umberto, Wyart, Matthieu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Probing the Latent Hierarchical Structure of Data via Diffusion Models
by: Sclocchi, Antonio, et al.
Published: (2024)
by: Sclocchi, Antonio, et al.
Published: (2024)
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures
by: Cagnetta, Francesco, et al.
Published: (2025)
by: Cagnetta, Francesco, et al.
Published: (2025)
On the different regimes of Stochastic Gradient Descent
by: Sclocchi, Antonio, et al.
Published: (2023)
by: Sclocchi, Antonio, et al.
Published: (2023)
Learning curves theory for hierarchically compositional data with power-law distributed features
by: Cagnetta, Francesco, et al.
Published: (2025)
by: Cagnetta, Francesco, et al.
Published: (2025)
A Phase Transition in Diffusion Models Reveals the Hierarchical Nature of Data
by: Sclocchi, Antonio, et al.
Published: (2024)
by: Sclocchi, Antonio, et al.
Published: (2024)
Deep networks learn to parse uniform-depth context-free languages from local statistics
by: Parley, Jack T., et al.
Published: (2026)
by: Parley, Jack T., et al.
Published: (2026)
Towards a theory of how the structure of language is acquired by deep neural networks
by: Cagnetta, Francesco, et al.
Published: (2024)
by: Cagnetta, Francesco, et al.
Published: (2024)
Sampling Data with Chains of Forward-Backward Diffusion Steps
by: Kang, Hyunmo, et al.
Published: (2026)
by: Kang, Hyunmo, et al.
Published: (2026)
On the Emergence of Linear Analogies in Word Embeddings
by: Korchinski, Daniel J., et al.
Published: (2025)
by: Korchinski, Daniel J., et al.
Published: (2025)
Symmetry in language statistics shapes the geometry of model representations
by: Karkada, Dhruva, et al.
Published: (2026)
by: Karkada, Dhruva, et al.
Published: (2026)
Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer
by: Bordelon, Blake, et al.
Published: (2025)
by: Bordelon, Blake, et al.
Published: (2025)
Deep Learning as Neural Low-Degree Filtering: A Spectral Theory of Hierarchical Feature Learning
by: Dandi, Yatin, et al.
Published: (2026)
by: Dandi, Yatin, et al.
Published: (2026)
Saddle Hierarchy in Dense Associative Memory
by: Thériault, Robin, et al.
Published: (2025)
by: Thériault, Robin, et al.
Published: (2025)
Theory of Optimal Learning Rate Schedules and Scaling Laws for a Random Feature Model
by: Bordelon, Blake, et al.
Published: (2026)
by: Bordelon, Blake, et al.
Published: (2026)
Statistical Physics of Deep Neural Networks: Generalization Capability, Beyond the Infinite Width, and Feature Learning
by: Ariosto, Sebastiano
Published: (2025)
by: Ariosto, Sebastiano
Published: (2025)
Asymmetric Scaling Laws from Sparse Features
by: Sous, John, et al.
Published: (2026)
by: Sous, John, et al.
Published: (2026)
A Random-Matrix Criterion for Initializing Gated Recurrent Neural Networks
by: Fioratti, Tommaso, et al.
Published: (2026)
by: Fioratti, Tommaso, et al.
Published: (2026)
The Interplay of Data Structure and Imbalance in the Learning Dynamics of Diffusion Models
by: Nicoletti, Flavio, et al.
Published: (2026)
by: Nicoletti, Flavio, et al.
Published: (2026)
A Theory of Saddle Escape in Deep Nonlinear Networks
by: Rawal, Divit, et al.
Published: (2026)
by: Rawal, Divit, et al.
Published: (2026)
Initial Guessing Bias: How Untrained Networks Favor Some Classes
by: Francazi, Emanuele, et al.
Published: (2023)
by: Francazi, Emanuele, et al.
Published: (2023)
Asymptotics of Learning with Deep Structured (Random) Features
by: Schröder, Dominik, et al.
Published: (2024)
by: Schröder, Dominik, et al.
Published: (2024)
Disordered Dynamics in High Dimensions: Connections to Random Matrices and Machine Learning
by: Bordelon, Blake, et al.
Published: (2026)
by: Bordelon, Blake, et al.
Published: (2026)
Exploring the Energy Landscape of RBMs: Reciprocal Space Insights into Bosons, Hierarchical Learning and Symmetry Breaking
by: Toledo-Marin, J. Quetzalcóatl, et al.
Published: (2025)
by: Toledo-Marin, J. Quetzalcóatl, et al.
Published: (2025)
Modeling Structured Data Learning with Restricted Boltzmann Machines in the Teacher-Student Setting
by: Thériault, Robin, et al.
Published: (2024)
by: Thériault, Robin, et al.
Published: (2024)
How Feature Learning Can Improve Neural Scaling Laws
by: Bordelon, Blake, et al.
Published: (2024)
by: Bordelon, Blake, et al.
Published: (2024)
EB-RANSAC: Random Sample Consensus based on Energy-Based Model
by: Yasuda, Muneki, et al.
Published: (2026)
by: Yasuda, Muneki, et al.
Published: (2026)
Sparse Interactions Reshape Stability in Random Lotka-Volterra Dynamics
by: Tarabolo, Mattia, et al.
Published: (2025)
by: Tarabolo, Mattia, et al.
Published: (2025)
The Training Process of Many Deep Networks Explores the Same Low-Dimensional Manifold
by: Mao, Jialin, et al.
Published: (2023)
by: Mao, Jialin, et al.
Published: (2023)
Small Singular Values Matter: A Random Matrix Analysis of Transformer Models
by: Staats, Max, et al.
Published: (2024)
by: Staats, Max, et al.
Published: (2024)
Transfer Learning in Infinite Width Feature Learning Networks
by: Lauditi, Clarissa, et al.
Published: (2025)
by: Lauditi, Clarissa, et al.
Published: (2025)
Kernel Renormalization in Bayesian Deep Neural Networks: the Equivalent Wishart Ansatz in the Proportional Regime
by: Baglioni, Paolo, et al.
Published: (2026)
by: Baglioni, Paolo, et al.
Published: (2026)
Random features and polynomial rules
by: Aguirre-López, Fabián, et al.
Published: (2024)
by: Aguirre-López, Fabián, et al.
Published: (2024)
Random Features Hopfield Networks generalize retrieval to previously unseen examples
by: Kalaj, Silvio, et al.
Published: (2024)
by: Kalaj, Silvio, et al.
Published: (2024)
The Rules-and-Facts Model for Simultaneous Generalization and Memorization in Neural Networks
by: Farné, Gabriele, et al.
Published: (2026)
by: Farné, Gabriele, et al.
Published: (2026)
(How) Can Transformers Predict Pseudo-Random Numbers?
by: Tao, Tao, et al.
Published: (2025)
by: Tao, Tao, et al.
Published: (2025)
Asymptotics of SGD in Sequence-Single Index Models and Single-Layer Attention Networks
by: Arnaboldi, Luca, et al.
Published: (2025)
by: Arnaboldi, Luca, et al.
Published: (2025)
Sparse chaos in cortical circuits
by: Engelken, Rainer, et al.
Published: (2024)
by: Engelken, Rainer, et al.
Published: (2024)
STEM Diffraction Pattern Analysis with Deep Learning Networks
by: Wissel, Sebastian, et al.
Published: (2025)
by: Wissel, Sebastian, et al.
Published: (2025)
Dynamical Learning in Deep Asymmetric Recurrent Neural Networks
by: Badalotti, Davide, et al.
Published: (2025)
by: Badalotti, Davide, et al.
Published: (2025)
No Free Lunch From Random Feature Ensembles: Scaling Laws and Near-Optimality Conditions
by: Ruben, Benjamin S., et al.
Published: (2024)
by: Ruben, Benjamin S., et al.
Published: (2024)
Similar Items
-
Probing the Latent Hierarchical Structure of Data via Diffusion Models
by: Sclocchi, Antonio, et al.
Published: (2024) -
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures
by: Cagnetta, Francesco, et al.
Published: (2025) -
On the different regimes of Stochastic Gradient Descent
by: Sclocchi, Antonio, et al.
Published: (2023) -
Learning curves theory for hierarchically compositional data with power-law distributed features
by: Cagnetta, Francesco, et al.
Published: (2025) -
A Phase Transition in Diffusion Models Reveals the Hierarchical Nature of Data
by: Sclocchi, Antonio, et al.
Published: (2024)