Position: Solve Layerwise Linear Models First to Understand Neural Dynamical Phenomena (Neural Collapse, Emergence, Lazy/Rich Regime, and Grokking)
Fuente:
arXiv
Saved in:
| Main Authors: | Nam, Yoonsoo, Lee, Seok Hyeong, Domine, Clementine C J, Park, Yeachan, London, Charles, Choi, Wonyl, Goring, Niclas, Lee, Seungjai |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cotype zeta functions enumerating subalgebras of $R$-algebras
by: Lee, Seok Hyeong, et al.
Published: (2025)
by: Lee, Seok Hyeong, et al.
Published: (2025)
A simple mean field model of feature learning
by: Göring, Niclas, et al.
Published: (2025)
by: Göring, Niclas, et al.
Published: (2025)
Decoupling Dynamical Richness from Representation Learning: Towards Practical Measurement
by: Nam, Yoonsoo, et al.
Published: (2024)
by: Nam, Yoonsoo, et al.
Published: (2024)
Zeta functions of quadratic lattices of a hyperbolic plane
by: Kim, Daejun, et al.
Published: (2025)
by: Kim, Daejun, et al.
Published: (2025)
Zeta functions of quadratic lattices of a hyperbolic plane
by: Daejun Kim, et al.
Published: (2025)
by: Daejun Kim, et al.
Published: (2025)
Zeta functions enumerating subforms of quadratic forms
by: Kim, Daejun, et al.
Published: (2024)
by: Kim, Daejun, et al.
Published: (2024)
Feature learning is decoupled from generalization in high capacity neural networks
by: Göring, Niclas Alexander, et al.
Published: (2025)
by: Göring, Niclas Alexander, et al.
Published: (2025)
From Lazy to Rich: Exact Learning Dynamics in Deep Linear Networks
by: Dominé, Clémentine C. J., et al.
Published: (2024)
by: Dominé, Clémentine C. J., et al.
Published: (2024)
Grokking as the Transition from Lazy to Rich Training Dynamics
by: Kumar, Tanishq, et al.
Published: (2023)
by: Kumar, Tanishq, et al.
Published: (2023)
Explaining Grokking and Information Bottleneck through Neural Collapse Emergence
by: Sakamoto, Keitaro, et al.
Published: (2025)
by: Sakamoto, Keitaro, et al.
Published: (2025)
Characterising the Inductive Biases of Neural Networks on Boolean Data
by: Mingard, Chris, et al.
Published: (2025)
by: Mingard, Chris, et al.
Published: (2025)
An exactly solvable model for emergence and scaling laws in the multitask sparse parity problem
by: Nam, Yoonsoo, et al.
Published: (2024)
by: Nam, Yoonsoo, et al.
Published: (2024)
Zeta functions of $\mathbb{F}_p$-Lie algebras and finite $p$-groups
by: Lee, Seungjai
Published: (2020)
by: Lee, Seungjai
Published: (2020)
Zeta functions of solvable Lie algebras over finite fields -- with calculations in detail
by: Lee, Seungjai
Published: (2026)
by: Lee, Seungjai
Published: (2026)
Acceleration of Grokking in Learning Arithmetic Operations via Kolmogorov-Arnold Representation
by: Park, Yeachan, et al.
Published: (2024)
by: Park, Yeachan, et al.
Published: (2024)
Flatness is Necessary, Neural Collapse is Not: Rethinking Generalization via Grokking
by: Han, Ting, et al.
Published: (2025)
by: Han, Ting, et al.
Published: (2025)
Expressive Power of Floating-Point Neural Networks with Arbitrary Reduction Orders and Inexact Activation Implementations
by: Park, Yeachan, et al.
Published: (2026)
by: Park, Yeachan, et al.
Published: (2026)
Mixed Dynamics In Linear Networks: Unifying the Lazy and Active Regimes
by: Tu, Zhenfeng, et al.
Published: (2024)
by: Tu, Zhenfeng, et al.
Published: (2024)
On Defining Neural Averaging
by: Lee, Su Hyeong, et al.
Published: (2025)
by: Lee, Su Hyeong, et al.
Published: (2025)
Floating-Point Neural Networks Are Provably Robust Universal Approximators
by: Hwang, Geonho, et al.
Published: (2025)
by: Hwang, Geonho, et al.
Published: (2025)
Linear Half-Space Problems in Kinetic Theory: Abstract Formulation and Regime Transitions
by: Bernhoff, Niclas
Published: (2022)
by: Bernhoff, Niclas
Published: (2022)
Semantic-Aware Gaussian Process Calibration with Structured Layerwise Kernels for Deep Neural Networks
by: Lee, Kyung-hwan, et al.
Published: (2025)
by: Lee, Kyung-hwan, et al.
Published: (2025)
Layerwise Change of Knowledge in Neural Networks
by: Cheng, Xu, et al.
Published: (2024)
by: Cheng, Xu, et al.
Published: (2024)
Optimal Representation Size: High-Dimensional Analysis of Pretraining and Linear Probing
by: Njaradi, Valentina, et al.
Published: (2026)
by: Njaradi, Valentina, et al.
Published: (2026)
Exploiting the equivalence between quantum neural networks and perceptrons
by: Mingard, Chris, et al.
Published: (2024)
by: Mingard, Chris, et al.
Published: (2024)
NeuralGrok: Accelerate Grokking by Neural Gradient Transformation
by: Zhou, Xinyu, et al.
Published: (2025)
by: Zhou, Xinyu, et al.
Published: (2025)
Lifting problem for universal quadratic forms over totally real cubic number fields
by: Kim, Daejun, et al.
Published: (2023)
by: Kim, Daejun, et al.
Published: (2023)
On Expressive Power of Quantized Neural Networks under Fixed-Point Arithmetic
by: Park, Yeachan, et al.
Published: (2024)
by: Park, Yeachan, et al.
Published: (2024)
Grokking as Dimensional Phase Transition in Neural Networks
by: Wang, Ping
Published: (2026)
by: Wang, Ping
Published: (2026)
CLIPSE -- a minimalistic CLIP-based image search engine for research
by: Göring, Steve
Published: (2025)
by: Göring, Steve
Published: (2025)
Grokking at the Edge of Linear Separability
by: Beck, Alon, et al.
Published: (2024)
by: Beck, Alon, et al.
Published: (2024)
Out-of-Domain Generalization in Dynamical Systems Reconstruction
by: Göring, Niclas, et al.
Published: (2024)
by: Göring, Niclas, et al.
Published: (2024)
Probabilistic Modeling of Latent Agentic Substructures in Deep Neural Networks
by: Lee, Su Hyeong, et al.
Published: (2025)
by: Lee, Su Hyeong, et al.
Published: (2025)
Neural Artistic Style and Color Transfer Using Deep Learning
by: London, Justin
Published: (2025)
by: London, Justin
Published: (2025)
Dominating vs. Dominated: Generative Collapse in Diffusion Models
by: Jeong, Hayeon, et al.
Published: (2025)
by: Jeong, Hayeon, et al.
Published: (2025)
Critical Phenomena in Gravitational Collapse
by: Gundlach, Carsten, et al.
Published: (2025)
by: Gundlach, Carsten, et al.
Published: (2025)
BALI: Learning Neural Networks via Bayesian Layerwise Inference
by: Kurle, Richard, et al.
Published: (2024)
by: Kurle, Richard, et al.
Published: (2024)
Absence of Closed-Form Descriptions for Gradient Flow in Two-Layer Narrow Networks
by: Park, Yeachan
Published: (2024)
by: Park, Yeachan
Published: (2024)
Grokfast: Accelerated Grokking by Amplifying Slow Gradients
by: Lee, Jaerin, et al.
Published: (2024)
by: Lee, Jaerin, et al.
Published: (2024)
Grokking and Generalization Collapse: Insights from \texttt{HTSR} theory
by: Prakash, Hari K., et al.
Published: (2025)
by: Prakash, Hari K., et al.
Published: (2025)
Similar Items
-
Cotype zeta functions enumerating subalgebras of $R$-algebras
by: Lee, Seok Hyeong, et al.
Published: (2025) -
A simple mean field model of feature learning
by: Göring, Niclas, et al.
Published: (2025) -
Decoupling Dynamical Richness from Representation Learning: Towards Practical Measurement
by: Nam, Yoonsoo, et al.
Published: (2024) -
Zeta functions of quadratic lattices of a hyperbolic plane
by: Kim, Daejun, et al.
Published: (2025) -
Zeta functions of quadratic lattices of a hyperbolic plane
by: Daejun Kim, et al.
Published: (2025)