Weight Conditioning for Smooth Optimization of Neural Networks

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Saratchandran, Hemanth, Wang, Thomas X., Lucey, Simon
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917336618893312
author Saratchandran, Hemanth
Wang, Thomas X.
Lucey, Simon
author_facet Saratchandran, Hemanth
Wang, Thomas X.
Lucey, Simon
contents In this article, we introduce a novel normalization technique for neural network weight matrices, which we term weight conditioning. This approach aims to narrow the gap between the smallest and largest singular values of the weight matrices, resulting in better-conditioned matrices. The inspiration for this technique partially derives from numerical linear algebra, where well-conditioned matrices are known to facilitate stronger convergence results for iterative solvers. We provide a theoretical foundation demonstrating that our normalization technique smoothens the loss landscape, thereby enhancing convergence of stochastic gradient descent algorithms. Empirically, we validate our normalization across various neural network architectures, including Convolutional Neural Networks (CNNs), Vision Transformers (ViT), Neural Radiance Fields (NeRF), and 3D shape modeling. Our findings indicate that our normalization method is not only competitive but also outperforms existing weight normalization techniques from the literature.
format Preprint
id arxiv_https___arxiv_org_abs_2409_03424
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Weight Conditioning for Smooth Optimization of Neural Networks
Saratchandran, Hemanth
Wang, Thomas X.
Lucey, Simon
Computer Vision and Pattern Recognition
In this article, we introduce a novel normalization technique for neural network weight matrices, which we term weight conditioning. This approach aims to narrow the gap between the smallest and largest singular values of the weight matrices, resulting in better-conditioned matrices. The inspiration for this technique partially derives from numerical linear algebra, where well-conditioned matrices are known to facilitate stronger convergence results for iterative solvers. We provide a theoretical foundation demonstrating that our normalization technique smoothens the loss landscape, thereby enhancing convergence of stochastic gradient descent algorithms. Empirically, we validate our normalization across various neural network architectures, including Convolutional Neural Networks (CNNs), Vision Transformers (ViT), Neural Radiance Fields (NeRF), and 3D shape modeling. Our findings indicate that our normalization method is not only competitive but also outperforms existing weight normalization techniques from the literature.
title Weight Conditioning for Smooth Optimization of Neural Networks
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2409.03424