From Lazy to Rich: Exact Learning Dynamics in Deep Linear Networks
Fuente:
arXiv
Saved in:
| Main Authors: | Dominé, Clémentine C. J., Anguita, Nicolas, Proca, Alexandra M., Braun, Lukas, Kunin, Daniel, Mediano, Pedro A. M., Saxe, Andrew M. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Theory of How Pretraining Shapes Inductive Bias in Fine-Tuning
by: Anguita, Nicolas, et al.
Published: (2026)
by: Anguita, Nicolas, et al.
Published: (2026)
Get rich quick: exact solutions reveal how unbalanced initializations promote rapid feature learning
by: Kunin, Daniel, et al.
Published: (2024)
by: Kunin, Daniel, et al.
Published: (2024)
Optimal Representation Size: High-Dimensional Analysis of Pretraining and Linear Probing
by: Njaradi, Valentina, et al.
Published: (2026)
by: Njaradi, Valentina, et al.
Published: (2026)
Position: Solve Layerwise Linear Models First to Understand Neural Dynamical Phenomena (Neural Collapse, Emergence, Lazy/Rich Regime, and Grokking)
by: Nam, Yoonsoo, et al.
Published: (2025)
by: Nam, Yoonsoo, et al.
Published: (2025)
Flexible task abstractions emerge in linear networks with fast and bounded units
by: Sandbrink, Kai, et al.
Published: (2024)
by: Sandbrink, Kai, et al.
Published: (2024)
A Theory of Initialisation's Impact on Specialisation
by: Jarvis, Devon, et al.
Published: (2025)
by: Jarvis, Devon, et al.
Published: (2025)
Understanding Unimodal Bias in Multimodal Deep Linear Networks
by: Zhang, Yedi, et al.
Published: (2023)
by: Zhang, Yedi, et al.
Published: (2023)
When Representations Align: Universality in Representation Learning Dynamics
by: van Rossem, Loek, et al.
Published: (2024)
by: van Rossem, Loek, et al.
Published: (2024)
Training Dynamics of In-Context Learning in Linear Attention
by: Zhang, Yedi, et al.
Published: (2025)
by: Zhang, Yedi, et al.
Published: (2025)
Grokking as the Transition from Lazy to Rich Training Dynamics
by: Kumar, Tanishq, et al.
Published: (2023)
by: Kumar, Tanishq, et al.
Published: (2023)
Mixed Dynamics In Linear Networks: Unifying the Lazy and Active Regimes
by: Tu, Zhenfeng, et al.
Published: (2024)
by: Tu, Zhenfeng, et al.
Published: (2024)
When Are Bias-Free ReLU Networks Effectively Linear Networks?
by: Zhang, Yedi, et al.
Published: (2024)
by: Zhang, Yedi, et al.
Published: (2024)
Meta-Learning Strategies through Value Maximization in Neural Networks
by: Carrasco-Davis, Rodrigo, et al.
Published: (2023)
by: Carrasco-Davis, Rodrigo, et al.
Published: (2023)
Algorithm Development in Neural Networks: Insights from the Streaming Parity Task
by: van Rossem, Loek, et al.
Published: (2025)
by: van Rossem, Loek, et al.
Published: (2025)
Softmax $\geq$ Linear: Transformers may learn to classify in-context by kernel gradient descent
by: Dragutinović, Sara, et al.
Published: (2025)
by: Dragutinović, Sara, et al.
Published: (2025)
Distinct Computations Emerge From Compositional Curricula in In-Context Learning
by: Lee, Jin Hwa, et al.
Published: (2025)
by: Lee, Jin Hwa, et al.
Published: (2025)
Make Haste Slowly: A Theory of Emergent Structured Mixed Selectivity in Feature Learning ReLU Networks
by: Jarvis, Devon, et al.
Published: (2025)
by: Jarvis, Devon, et al.
Published: (2025)
Saddle-to-Saddle Dynamics Explains A Simplicity Bias Across Neural Network Architectures
by: Zhang, Yedi, et al.
Published: (2025)
by: Zhang, Yedi, et al.
Published: (2025)
Feature Learning beyond the Lazy-Rich Dichotomy: Insights from Representational Geometry
by: Chou, Chi-Ning, et al.
Published: (2025)
by: Chou, Chi-Ning, et al.
Published: (2025)
The Omniscient, yet Lazy, Investor
by: Halkiewicz, Stanisław M. S.
Published: (2025)
by: Halkiewicz, Stanisław M. S.
Published: (2025)
LazyDiT: Lazy Learning for the Acceleration of Diffusion Transformers
by: Shen, Xuan, et al.
Published: (2024)
by: Shen, Xuan, et al.
Published: (2024)
¿Confidencialidad, anonimato?: las otras promesas de la investigación
by: Verónica Anguita M.
Published: (2011)
by: Verónica Anguita M.
Published: (2011)
Lazy Linearity for a Core Functional Language
by: Mesquita, Rodrigo, et al.
Published: (2025)
by: Mesquita, Rodrigo, et al.
Published: (2025)
A Logarithmic Decomposition and a Signed Measure Space for Entropy
by: Down, Keenan J. A., et al.
Published: (2024)
by: Down, Keenan J. A., et al.
Published: (2024)
Algebraic Representations of Entropy and Fixed-Parity Information Quantities
by: Down, Keenan J. A., et al.
Published: (2024)
by: Down, Keenan J. A., et al.
Published: (2024)
Synergistic Motifs in Gaussian Systems
by: Caprioglio, Enrico, et al.
Published: (2025)
by: Caprioglio, Enrico, et al.
Published: (2025)
APUNTES SOBRE EL RAP POLÍTICO BOLIVIANO
by: Johana Kunin
Published: (2009)
by: Johana Kunin
Published: (2009)
Percepción social del riesgo y dinâmicas de género en la producción agrícola basada en plaguicidas en la pampa húmeda Argentina
by: Johana Kunin
Published: (2020)
by: Johana Kunin
Published: (2020)
Reseña de "Ciberculturas juveniles. Los jóvenes, sus prácticas y sus representaciones en la era de Internet" de Marcelo Urresti (edit.)
by: Johana Kunin
Published: (2008)
by: Johana Kunin
Published: (2008)
Pariendo madres: talleres de parto y de crianza en un distrito rural bonaerense
by: Johana Kunin
Published: (2021)
by: Johana Kunin
Published: (2021)
Los "medio putos": masculinidades subalternas y dinámicas de género alternativas en la rural Pampa húmeda argentina (2014-2017)
by: Johana Kunin
Published: (2021)
by: Johana Kunin
Published: (2021)
Cuando los paradigmas agroecológicos y del parto respetado no encuentran las respuestas esperadas
by: Johana Kunin
Published: (2022)
by: Johana Kunin
Published: (2022)
“Acá se sabe si la gente se aísla”: (anti)anonimato, cuidado y poder en localidades medianas y pequeñas en tiempos de COVID-19
by: Johana Kunin
Published: (2021)
by: Johana Kunin
Published: (2021)
Sequential Group Composition: A Window into the Mechanics of Deep Learning
by: Marchetti, Giovanni Luca, et al.
Published: (2026)
by: Marchetti, Giovanni Luca, et al.
Published: (2026)
From Data Statistics to Feature Geometry: How Correlations Shape Superposition
by: Prieto, Lucas, et al.
Published: (2026)
by: Prieto, Lucas, et al.
Published: (2026)
Nonlinear dynamics of localization in neural receptive fields
by: Lufkin, Leon, et al.
Published: (2025)
by: Lufkin, Leon, et al.
Published: (2025)
Hipersonolência diurna e variáveis polissonográficas em doentes com síndroma de apneia do sono
by: O Mediano
Published: (2007)
by: O Mediano
Published: (2007)
Designing Walrus: Relational Programming with Rich Types, On-Demand Laziness, and Structured Traces
by: Cuéllar, Santiago, et al.
Published: (2025)
by: Cuéllar, Santiago, et al.
Published: (2025)
Lazy-DaSH: Lazy Approach for Hypergraph-based Multi-robot Task and Motion Planning
by: Lee, Seongwon, et al.
Published: (2025)
by: Lee, Seongwon, et al.
Published: (2025)
Directional Confusions Reveal Divergent Inductive Biases Through Rate-Distortion Geometry in Human and Machine Vision
by: Caglar, Leyla Roksan, et al.
Published: (2026)
by: Caglar, Leyla Roksan, et al.
Published: (2026)
Similar Items
-
A Theory of How Pretraining Shapes Inductive Bias in Fine-Tuning
by: Anguita, Nicolas, et al.
Published: (2026) -
Get rich quick: exact solutions reveal how unbalanced initializations promote rapid feature learning
by: Kunin, Daniel, et al.
Published: (2024) -
Optimal Representation Size: High-Dimensional Analysis of Pretraining and Linear Probing
by: Njaradi, Valentina, et al.
Published: (2026) -
Position: Solve Layerwise Linear Models First to Understand Neural Dynamical Phenomena (Neural Collapse, Emergence, Lazy/Rich Regime, and Grokking)
by: Nam, Yoonsoo, et al.
Published: (2025) -
Flexible task abstractions emerge in linear networks with fast and bounded units
by: Sandbrink, Kai, et al.
Published: (2024)