IDInit: A Universal and Stable Initialization Method for Neural Network Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pan, Yu, Wang, Chaozheng, Wu, Zekai, Wang, Qifan, Zhang, Min, Xu, Zenglin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913727065882624
author Pan, Yu
Wang, Chaozheng
Wu, Zekai
Wang, Qifan
Zhang, Min
Xu, Zenglin
author_facet Pan, Yu
Wang, Chaozheng
Wu, Zekai
Wang, Qifan
Zhang, Min
Xu, Zenglin
contents Deep neural networks have achieved remarkable accomplishments in practice. The success of these networks hinges on effective initialization methods, which are vital for ensuring stable and rapid convergence during training. Recently, initialization methods that maintain identity transition within layers have shown good efficiency in network training. These techniques (e.g., Fixup) set specific weights to zero to achieve identity control. However, settings of remaining weight (e.g., Fixup uses random values to initialize non-zero weights) will affect the inductive bias that is achieved only by a zero weight, which may be harmful to training. Addressing this concern, we introduce fully identical initialization (IDInit), a novel method that preserves identity in both the main and sub-stem layers of residual networks. IDInit employs a padded identity-like matrix to overcome rank constraints in non-square weight matrices. Furthermore, we show the convergence problem of an identity matrix can be solved by stochastic gradient descent. Additionally, we enhance the universality of IDInit by processing higher-order weights and addressing dead neuron problems. IDInit is a straightforward yet effective initialization method, with improved convergence, stability, and performance across various settings, including large-scale datasets and deep models.
format Preprint
id arxiv_https___arxiv_org_abs_2503_04626
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle IDInit: A Universal and Stable Initialization Method for Neural Network Training
Pan, Yu
Wang, Chaozheng
Wu, Zekai
Wang, Qifan
Zhang, Min
Xu, Zenglin
Machine Learning
Artificial Intelligence
Deep neural networks have achieved remarkable accomplishments in practice. The success of these networks hinges on effective initialization methods, which are vital for ensuring stable and rapid convergence during training. Recently, initialization methods that maintain identity transition within layers have shown good efficiency in network training. These techniques (e.g., Fixup) set specific weights to zero to achieve identity control. However, settings of remaining weight (e.g., Fixup uses random values to initialize non-zero weights) will affect the inductive bias that is achieved only by a zero weight, which may be harmful to training. Addressing this concern, we introduce fully identical initialization (IDInit), a novel method that preserves identity in both the main and sub-stem layers of residual networks. IDInit employs a padded identity-like matrix to overcome rank constraints in non-square weight matrices. Furthermore, we show the convergence problem of an identity matrix can be solved by stochastic gradient descent. Additionally, we enhance the universality of IDInit by processing higher-order weights and addressing dead neuron problems. IDInit is a straightforward yet effective initialization method, with improved convergence, stability, and performance across various settings, including large-scale datasets and deep models.
title IDInit: A Universal and Stable Initialization Method for Neural Network Training
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2503.04626