The Role of Symmetry in Optimizing Overparameterized Networks

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Sareen, Kusha, Pedramfar, Mohammad, Kaba, Sékou-Oumar, Shakerinava, Mehran, Ravanbakhsh, Siamak
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909025017266176
author Sareen, Kusha
Pedramfar, Mohammad
Kaba, Sékou-Oumar
Shakerinava, Mehran
Ravanbakhsh, Siamak
author_facet Sareen, Kusha
Pedramfar, Mohammad
Kaba, Sékou-Oumar
Shakerinava, Mehran
Ravanbakhsh, Siamak
contents Overparameterization is central to the success of deep learning, yet the mechanisms by which it improves optimization remain incompletely understood. We analyze weight-space symmetries in neural networks and show that overparameterization introduces additional symmetries that benefit optimization in two distinct ways. First, we prove that these symmetries act as a form of diagonal preconditioning on the Hessian, enabling the existence of better-conditioned minima within each equivalence class of functionally identical solutions. Second, we show that overparameterization increases the probability mass of global minima near typical initializations, making these favourable solutions more reachable. These results offer a potential link between loss landscape geometry and simplicity bias. Empirically, we observe wider networks have lower top eigenvalues, smaller condition numbers and faster convergence, matching our analysis. Our analysis provides a unified framework for understanding overparameterization and width growth as a geometric transformation of the loss landscape.
format Preprint
id arxiv_https___arxiv_org_abs_2604_25150
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle The Role of Symmetry in Optimizing Overparameterized Networks
Sareen, Kusha
Pedramfar, Mohammad
Kaba, Sékou-Oumar
Shakerinava, Mehran
Ravanbakhsh, Siamak
Machine Learning
Artificial Intelligence
Overparameterization is central to the success of deep learning, yet the mechanisms by which it improves optimization remain incompletely understood. We analyze weight-space symmetries in neural networks and show that overparameterization introduces additional symmetries that benefit optimization in two distinct ways. First, we prove that these symmetries act as a form of diagonal preconditioning on the Hessian, enabling the existence of better-conditioned minima within each equivalence class of functionally identical solutions. Second, we show that overparameterization increases the probability mass of global minima near typical initializations, making these favourable solutions more reachable. These results offer a potential link between loss landscape geometry and simplicity bias. Empirically, we observe wider networks have lower top eigenvalues, smaller condition numbers and faster convergence, matching our analysis. Our analysis provides a unified framework for understanding overparameterization and width growth as a geometric transformation of the loss landscape.
title The Role of Symmetry in Optimizing Overparameterized Networks
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2604.25150