Where You Place the Norm Matters: From Prejudiced to Neutral Initializations
Fuente:
arXiv
Saved in:
| Main Authors: | Francazi, Emanuele, Pinto, Francesco, Lucchi, Aurelien, Baity-Jesi, Marco |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Initial Guessing Bias: How Untrained Networks Favor Some Classes
by: Francazi, Emanuele, et al.
Published: (2023)
by: Francazi, Emanuele, et al.
Published: (2023)
Class Imbalance in Anomaly Detection: Learning from an Exactly Solvable Model
by: Pezzicoli, F. S., et al.
Published: (2025)
by: Pezzicoli, F. S., et al.
Published: (2025)
A Theoretical Analysis of the Learning Dynamics under Class Imbalance
by: Francazi, Emanuele, et al.
Published: (2022)
by: Francazi, Emanuele, et al.
Published: (2022)
A Random-Matrix Criterion for Initializing Gated Recurrent Neural Networks
by: Fioratti, Tommaso, et al.
Published: (2026)
by: Fioratti, Tommaso, et al.
Published: (2026)
When Bias Meets Trainability: Connecting Theories of Initialization
by: Bassi, Alberto, et al.
Published: (2025)
by: Bassi, Alberto, et al.
Published: (2025)
Dataset-Free Weight-Initialization on Restricted Boltzmann Machine
by: Yasuda, Muneki, et al.
Published: (2024)
by: Yasuda, Muneki, et al.
Published: (2024)
When Can You Get Away with Low Memory Adam?
by: Kalra, Dayal Singh, et al.
Published: (2025)
by: Kalra, Dayal Singh, et al.
Published: (2025)
Restoring balance: principled under/oversampling of data for optimal classification
by: Loffredo, Emanuele, et al.
Published: (2024)
by: Loffredo, Emanuele, et al.
Published: (2024)
Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer
by: Bordelon, Blake, et al.
Published: (2025)
by: Bordelon, Blake, et al.
Published: (2025)
Fundamental limits of learning in sequence multi-index models and deep attention networks: High-dimensional asymptotics and sharp thresholds
by: Troiani, Emanuele, et al.
Published: (2025)
by: Troiani, Emanuele, et al.
Published: (2025)
Bilinear Sequence Regression: A Model for Learning from Long Sequences of High-dimensional Tokens
by: Erba, Vittorio, et al.
Published: (2024)
by: Erba, Vittorio, et al.
Published: (2024)
Small Singular Values Matter: A Random Matrix Analysis of Transformer Models
by: Staats, Max, et al.
Published: (2024)
by: Staats, Max, et al.
Published: (2024)
Theory of Optimal Learning Rate Schedules and Scaling Laws for a Random Feature Model
by: Bordelon, Blake, et al.
Published: (2026)
by: Bordelon, Blake, et al.
Published: (2026)
Self-attention as an attractor network: transient memories without backpropagation
by: D'Amico, Francesco, et al.
Published: (2024)
by: D'Amico, Francesco, et al.
Published: (2024)
Learning curves theory for hierarchically compositional data with power-law distributed features
by: Cagnetta, Francesco, et al.
Published: (2025)
by: Cagnetta, Francesco, et al.
Published: (2025)
Modeling Structured Data Learning with Restricted Boltzmann Machines in the Teacher-Student Setting
by: Thériault, Robin, et al.
Published: (2024)
by: Thériault, Robin, et al.
Published: (2024)
A solvable high-dimensional model where nonlinear autoencoders learn structure invisible to PCA while test loss misaligns with generalization
by: Mendes, Vicente Conde, et al.
Published: (2026)
by: Mendes, Vicente Conde, et al.
Published: (2026)
Implicit bias produces neural scaling laws in learning curves, from perceptrons to deep networks
by: D'Amico, Francesco, et al.
Published: (2025)
by: D'Amico, Francesco, et al.
Published: (2025)
Training neural networks with structured noise improves classification and generalization
by: Benedetti, Marco, et al.
Published: (2023)
by: Benedetti, Marco, et al.
Published: (2023)
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures
by: Cagnetta, Francesco, et al.
Published: (2025)
by: Cagnetta, Francesco, et al.
Published: (2025)
Regularization, early-stopping and dreaming: a Hopfield-like setup to address generalization and overfitting
by: Agliari, Elena, et al.
Published: (2023)
by: Agliari, Elena, et al.
Published: (2023)
The Copycat Perceptron: Smashing Barriers Through Collective Learning
by: Catania, Giovanni, et al.
Published: (2023)
by: Catania, Giovanni, et al.
Published: (2023)
The Symmetric Perceptron: a Teacher-Student Scenario
by: Catania, Giovanni, et al.
Published: (2026)
by: Catania, Giovanni, et al.
Published: (2026)
Theory of Speciation Transitions in Diffusion Models with General Class Structure
by: Achilli, Beatrice, et al.
Published: (2026)
by: Achilli, Beatrice, et al.
Published: (2026)
Towards Understanding Inductive Bias in Transformers: A View From Infinity
by: Lavie, Itay, et al.
Published: (2024)
by: Lavie, Itay, et al.
Published: (2024)
From Kernels to Features: A Multi-Scale Adaptive Theory of Feature Learning
by: Rubin, Noa, et al.
Published: (2025)
by: Rubin, Noa, et al.
Published: (2025)
No Free Lunch From Random Feature Ensembles: Scaling Laws and Near-Optimality Conditions
by: Ruben, Benjamin S., et al.
Published: (2024)
by: Ruben, Benjamin S., et al.
Published: (2024)
A theoretical framework for overfitting in energy-based modeling
by: Catania, Giovanni, et al.
Published: (2025)
by: Catania, Giovanni, et al.
Published: (2025)
On the role of non-linear latent features in bipartite generative neural networks
by: Bonnaire, Tony, et al.
Published: (2025)
by: Bonnaire, Tony, et al.
Published: (2025)
Explaining the effects of non-convergent sampling in the training of Energy-Based Models
by: Agoritsas, Elisabeth, et al.
Published: (2023)
by: Agoritsas, Elisabeth, et al.
Published: (2023)
Cascade of phase transitions in the training of Energy-based models
by: Bachtis, Dimitrios, et al.
Published: (2024)
by: Bachtis, Dimitrios, et al.
Published: (2024)
Benchmarking Graph Neural Networks in Solving Hard Constraint Satisfaction Problems
by: Skenderi, Geri, et al.
Published: (2026)
by: Skenderi, Geri, et al.
Published: (2026)
Inferring Higher-Order Couplings with Neural Networks
by: Decelle, Aurélien, et al.
Published: (2025)
by: Decelle, Aurélien, et al.
Published: (2025)
Fast training and sampling of Restricted Boltzmann Machines
by: Béreux, Nicolas, et al.
Published: (2024)
by: Béreux, Nicolas, et al.
Published: (2024)
Thermodynamics of bidirectional associative memories
by: Barra, Adriano, et al.
Published: (2022)
by: Barra, Adriano, et al.
Published: (2022)
Supervised Hebbian Learning
by: Alemanno, Francesco, et al.
Published: (2022)
by: Alemanno, Francesco, et al.
Published: (2022)
Inferring effective couplings with Restricted Boltzmann Machines
by: Decelle, Aurélien, et al.
Published: (2023)
by: Decelle, Aurélien, et al.
Published: (2023)
Exact full-RSB SAT/UNSAT transition in infinitely wide two-layer neural networks
by: Annesi, Brandon L., et al.
Published: (2024)
by: Annesi, Brandon L., et al.
Published: (2024)
Bayes optimal learning of attention-indexed models
by: Boncoraglio, Fabrizio, et al.
Published: (2025)
by: Boncoraglio, Fabrizio, et al.
Published: (2025)
The Nuclear Route: Sharp Asymptotics of ERM in Overparameterized Quadratic Networks
by: Erba, Vittorio, et al.
Published: (2025)
by: Erba, Vittorio, et al.
Published: (2025)
Similar Items
-
Initial Guessing Bias: How Untrained Networks Favor Some Classes
by: Francazi, Emanuele, et al.
Published: (2023) -
Class Imbalance in Anomaly Detection: Learning from an Exactly Solvable Model
by: Pezzicoli, F. S., et al.
Published: (2025) -
A Theoretical Analysis of the Learning Dynamics under Class Imbalance
by: Francazi, Emanuele, et al.
Published: (2022) -
A Random-Matrix Criterion for Initializing Gated Recurrent Neural Networks
by: Fioratti, Tommaso, et al.
Published: (2026) -
When Bias Meets Trainability: Connecting Theories of Initialization
by: Bassi, Alberto, et al.
Published: (2025)