Position: Solve Layerwise Linear Models First to Understand Neural Dynamical Phenomena (Neural Collapse, Emergence, Lazy/Rich Regime, and Grokking)

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Nam, Yoonsoo, Lee, Seok Hyeong, Domine, Clementine C J, Park, Yeachan, London, Charles, Choi, Wonyl, Goring, Niclas, Lee, Seungjai
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908379819016192
author Nam, Yoonsoo
Lee, Seok Hyeong
Domine, Clementine C J
Park, Yeachan
London, Charles
Choi, Wonyl
Goring, Niclas
Lee, Seungjai
author_facet Nam, Yoonsoo
Lee, Seok Hyeong
Domine, Clementine C J
Park, Yeachan
London, Charles
Choi, Wonyl
Goring, Niclas
Lee, Seungjai
contents In physics, complex systems are often simplified into minimal, solvable models that retain only the core principles. In machine learning, layerwise linear models (e.g., linear neural networks) act as simplified representations of neural network dynamics. These models follow the dynamical feedback principle, which describes how layers mutually govern and amplify each other's evolution. This principle extends beyond the simplified models, successfully explaining a wide range of dynamical phenomena in deep neural networks, including neural collapse, emergence, lazy and rich regimes, and grokking. In this position paper, we call for the use of layerwise linear models retaining the core principles of neural dynamical phenomena to accelerate the science of deep learning.
format Preprint
id arxiv_https___arxiv_org_abs_2502_21009
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Position: Solve Layerwise Linear Models First to Understand Neural Dynamical Phenomena (Neural Collapse, Emergence, Lazy/Rich Regime, and Grokking)
Nam, Yoonsoo
Lee, Seok Hyeong
Domine, Clementine C J
Park, Yeachan
London, Charles
Choi, Wonyl
Goring, Niclas
Lee, Seungjai
Machine Learning
Data Analysis, Statistics and Probability
In physics, complex systems are often simplified into minimal, solvable models that retain only the core principles. In machine learning, layerwise linear models (e.g., linear neural networks) act as simplified representations of neural network dynamics. These models follow the dynamical feedback principle, which describes how layers mutually govern and amplify each other's evolution. This principle extends beyond the simplified models, successfully explaining a wide range of dynamical phenomena in deep neural networks, including neural collapse, emergence, lazy and rich regimes, and grokking. In this position paper, we call for the use of layerwise linear models retaining the core principles of neural dynamical phenomena to accelerate the science of deep learning.
title Position: Solve Layerwise Linear Models First to Understand Neural Dynamical Phenomena (Neural Collapse, Emergence, Lazy/Rich Regime, and Grokking)
topic Machine Learning
Data Analysis, Statistics and Probability
url https://arxiv.org/abs/2502.21009