A Theory of Initialisation's Impact on Specialisation
Fuente:
arXiv
Saved in:
| Main Authors: | Jarvis, Devon, Lee, Sebastian, Dominé, Clémentine Carla Juliette, Saxe, Andrew M, Mannelli, Stefano Sarao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The RL Perceptron: Generalisation Dynamics of Policy Learning in High Dimensions
by: Patel, Nishil, et al.
Published: (2023)
by: Patel, Nishil, et al.
Published: (2023)
Position: the Stochastic Parrot in the Coal Mine. Model Collapse is a Threat to Low-Resource Communities
by: Jarvis, Devon, et al.
Published: (2026)
by: Jarvis, Devon, et al.
Published: (2026)
Tilting the Odds at the Lottery: the Interplay of Overparameterisation and Curricula in Neural Networks
by: Mannelli, Stefano Sarao, et al.
Published: (2024)
by: Mannelli, Stefano Sarao, et al.
Published: (2024)
Why Do Animals Need Shaping? A Theory of Task Composition and Curriculum Learning
by: Lee, Jin Hwa, et al.
Published: (2024)
by: Lee, Jin Hwa, et al.
Published: (2024)
Make Haste Slowly: A Theory of Emergent Structured Mixed Selectivity in Feature Learning ReLU Networks
by: Jarvis, Devon, et al.
Published: (2025)
by: Jarvis, Devon, et al.
Published: (2025)
Optimal Protocols for Continual Learning via Statistical Physics and Control Theory
by: Mori, Francesco, et al.
Published: (2024)
by: Mori, Francesco, et al.
Published: (2024)
Optimal Representation Size: High-Dimensional Analysis of Pretraining and Linear Probing
by: Njaradi, Valentina, et al.
Published: (2026)
by: Njaradi, Valentina, et al.
Published: (2026)
A Theory of How Pretraining Shapes Inductive Bias in Fine-Tuning
by: Anguita, Nicolas, et al.
Published: (2026)
by: Anguita, Nicolas, et al.
Published: (2026)
Sharp description of local minima in the loss landscape of high-dimensional two-layer ReLU neural networks
by: Huang, Jie, et al.
Published: (2026)
by: Huang, Jie, et al.
Published: (2026)
On The Specialization of Neural Modules
by: Jarvis, Devon, et al.
Published: (2024)
by: Jarvis, Devon, et al.
Published: (2024)
Revisiting the Role of Relearning in Semantic Dementia
by: Jarvis, Devon, et al.
Published: (2025)
by: Jarvis, Devon, et al.
Published: (2025)
Bias in Motion: Theoretical Insights into the Dynamics of Bias in SGD Training
by: Jain, Anchit, et al.
Published: (2024)
by: Jain, Anchit, et al.
Published: (2024)
Bias-inducing geometries: an exactly solvable data model with fairness implications
by: Mannelli, Stefano Sarao, et al.
Published: (2022)
by: Mannelli, Stefano Sarao, et al.
Published: (2022)
Get rich quick: exact solutions reveal how unbalanced initializations promote rapid feature learning
by: Kunin, Daniel, et al.
Published: (2024)
by: Kunin, Daniel, et al.
Published: (2024)
The Interplay of Data Structure and Imbalance in the Learning Dynamics of Diffusion Models
by: Nicoletti, Flavio, et al.
Published: (2026)
by: Nicoletti, Flavio, et al.
Published: (2026)
From Lazy to Rich: Exact Learning Dynamics in Deep Linear Networks
by: Dominé, Clémentine C. J., et al.
Published: (2024)
by: Dominé, Clémentine C. J., et al.
Published: (2024)
Reducing Recurrent Competitive Epidemics via Dynamic Resource Allocation
by: Kalogeratos, Argyris, et al.
Published: (2020)
by: Kalogeratos, Argyris, et al.
Published: (2020)
Pruning at Initialisation through the lens of Graphon Limit: Convergence, Expressivity, and Generalisation
by: Pham, Hoang, et al.
Published: (2026)
by: Pham, Hoang, et al.
Published: (2026)
Neural Architecture Search: Two Constant Shared Weights Initialisations
by: Gracheva, Ekaterina
Published: (2023)
by: Gracheva, Ekaterina
Published: (2023)
Nonlinear dynamics of localization in neural receptive fields
by: Lufkin, Leon, et al.
Published: (2025)
by: Lufkin, Leon, et al.
Published: (2025)
Algorithm Development in Neural Networks: Insights from the Streaming Parity Task
by: van Rossem, Loek, et al.
Published: (2025)
by: van Rossem, Loek, et al.
Published: (2025)
When Representations Align: Universality in Representation Learning Dynamics
by: van Rossem, Loek, et al.
Published: (2024)
by: van Rossem, Loek, et al.
Published: (2024)
Approximate Gaussianity Beyond Initialisation in Neural Networks
by: Hirst, Edward, et al.
Published: (2025)
by: Hirst, Edward, et al.
Published: (2025)
Saddle-to-Saddle Dynamics Explains A Simplicity Bias Across Neural Network Architectures
by: Zhang, Yedi, et al.
Published: (2025)
by: Zhang, Yedi, et al.
Published: (2025)
Initialisation Determines the Basin: Efficient Codebook Optimisation for Extreme LLM Quantization
by: Kennedy, Ian W., et al.
Published: (2026)
by: Kennedy, Ian W., et al.
Published: (2026)
Softmax $\geq$ Linear: Transformers may learn to classify in-context by kernel gradient descent
by: Dragutinović, Sara, et al.
Published: (2025)
by: Dragutinović, Sara, et al.
Published: (2025)
Skill Machines: Temporal Logic Skill Composition in Reinforcement Learning
by: Tasse, Geraud Nangue, et al.
Published: (2022)
by: Tasse, Geraud Nangue, et al.
Published: (2022)
Initialisation and Network Effects in Decentralised Federated Learning
by: Badie-Modiri, Arash, et al.
Published: (2024)
by: Badie-Modiri, Arash, et al.
Published: (2024)
Understanding Unimodal Bias in Multimodal Deep Linear Networks
by: Zhang, Yedi, et al.
Published: (2023)
by: Zhang, Yedi, et al.
Published: (2023)
When Are Bias-Free ReLU Networks Effectively Linear Networks?
by: Zhang, Yedi, et al.
Published: (2024)
by: Zhang, Yedi, et al.
Published: (2024)
Distinct Computations Emerge From Compositional Curricula in In-Context Learning
by: Lee, Jin Hwa, et al.
Published: (2025)
by: Lee, Jin Hwa, et al.
Published: (2025)
Position: Solve Layerwise Linear Models First to Understand Neural Dynamical Phenomena (Neural Collapse, Emergence, Lazy/Rich Regime, and Grokking)
by: Nam, Yoonsoo, et al.
Published: (2025)
by: Nam, Yoonsoo, et al.
Published: (2025)
Early learning of the optimal constant solution in neural networks and humans
by: Rubruck, Jirko, et al.
Published: (2024)
by: Rubruck, Jirko, et al.
Published: (2024)
Training Dynamics of In-Context Learning in Linear Attention
by: Zhang, Yedi, et al.
Published: (2025)
by: Zhang, Yedi, et al.
Published: (2025)
Meta-Learning Strategies through Value Maximization in Neural Networks
by: Carrasco-Davis, Rodrigo, et al.
Published: (2023)
by: Carrasco-Davis, Rodrigo, et al.
Published: (2023)
What needs to go right for an induction head? A mechanistic study of in-context learning circuits and their formation
by: Singh, Aaditya K., et al.
Published: (2024)
by: Singh, Aaditya K., et al.
Published: (2024)
Specialised or Generic? Tokenization Choices for Radiology Language Models
by: Warr, Hermione, et al.
Published: (2025)
by: Warr, Hermione, et al.
Published: (2025)
Strategy Coopetition Explains the Emergence and Transience of In-Context Learning
by: Singh, Aaditya K., et al.
Published: (2025)
by: Singh, Aaditya K., et al.
Published: (2025)
Normalisation and Initialisation Strategies for Graph Neural Networks in Blockchain Anomaly Detection
by: Duy, Dang Sy, et al.
Published: (2026)
by: Duy, Dang Sy, et al.
Published: (2026)
Flexible task abstractions emerge in linear networks with fast and bounded units
by: Sandbrink, Kai, et al.
Published: (2024)
by: Sandbrink, Kai, et al.
Published: (2024)
Similar Items
-
The RL Perceptron: Generalisation Dynamics of Policy Learning in High Dimensions
by: Patel, Nishil, et al.
Published: (2023) -
Position: the Stochastic Parrot in the Coal Mine. Model Collapse is a Threat to Low-Resource Communities
by: Jarvis, Devon, et al.
Published: (2026) -
Tilting the Odds at the Lottery: the Interplay of Overparameterisation and Curricula in Neural Networks
by: Mannelli, Stefano Sarao, et al.
Published: (2024) -
Why Do Animals Need Shaping? A Theory of Task Composition and Curriculum Learning
by: Lee, Jin Hwa, et al.
Published: (2024) -
Make Haste Slowly: A Theory of Emergent Structured Mixed Selectivity in Feature Learning ReLU Networks
by: Jarvis, Devon, et al.
Published: (2025)