Optimal Condition for Initialization Variance in Deep Neural Networks: An SGD Dynamics Perspective
Fuente:
arXiv
Saved in:
| Main Authors: | Horii, Hiroshi, Has, Sothea |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Weight Initialization and Variance Dynamics in Deep Neural Networks and Large Language Models
by: Han, Yankun
Published: (2025)
by: Han, Yankun
Published: (2025)
SGD with Partial Hessian for Deep Neural Networks Optimization
by: Sun, Ying, et al.
Published: (2024)
by: Sun, Ying, et al.
Published: (2024)
From SGD to Spectra: A Theory of Neural Network Weight Dynamics
by: Olsen, Brian Richard, et al.
Published: (2025)
by: Olsen, Brian Richard, et al.
Published: (2025)
Exploring and Improving Initialization for Deep Graph Neural Networks: A Signal Propagation Perspective
by: Wang, Senmiao, et al.
Published: (2025)
by: Wang, Senmiao, et al.
Published: (2025)
SGD for Variational Inference: Tackling Unbounded Variance via Preconditioning and Dynamic Batching
by: Labarrière, Hippolyte, et al.
Published: (2026)
by: Labarrière, Hippolyte, et al.
Published: (2026)
Deep Neural Network Initialization with Sparsity Inducing Activations
by: Price, Ilan, et al.
Published: (2024)
by: Price, Ilan, et al.
Published: (2024)
Effects of Initialization Biases on Deep Neural Network Training Dynamics
by: Pellegrino, Nicholas, et al.
Published: (2025)
by: Pellegrino, Nicholas, et al.
Published: (2025)
VAMO: Efficient Zeroth-Order Variance Reduction for SGD with Faster Convergence
by: Chen, Jiahe, et al.
Published: (2025)
by: Chen, Jiahe, et al.
Published: (2025)
Fair CoVariance Neural Networks
by: Cavallo, Andrea, et al.
Published: (2024)
by: Cavallo, Andrea, et al.
Published: (2024)
The Butterfly Effect: Neural Network Training Trajectories Are Highly Sensitive to Initial Conditions
by: Kwok, Devin, et al.
Published: (2025)
by: Kwok, Devin, et al.
Published: (2025)
SGD and Weight Decay Secretly Minimize the Rank of Your Neural Network
by: Galanti, Tomer, et al.
Published: (2022)
by: Galanti, Tomer, et al.
Published: (2022)
Deep Operator Networks for Surrogate Modeling of Cyclic Adsorption Processes with Varying Initial Conditions
by: Ceccanti, Beatrice, et al.
Published: (2026)
by: Ceccanti, Beatrice, et al.
Published: (2026)
Optimal Initialization in Depth: Lyapunov Initialization and Limit Theorems for Deep Leaky ReLU Networks
by: Kogler, Constantin, et al.
Published: (2026)
by: Kogler, Constantin, et al.
Published: (2026)
DP-SGD Without Clipping: The Lipschitz Neural Network Way
by: Bethune, Louis, et al.
Published: (2023)
by: Bethune, Louis, et al.
Published: (2023)
Optimized Weight Initialization on the Stiefel Manifold for Deep ReLU Neural Networks
by: Lee, Hyungu, et al.
Published: (2025)
by: Lee, Hyungu, et al.
Published: (2025)
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training
by: Sadrtdinov, Ildus, et al.
Published: (2025)
by: Sadrtdinov, Ildus, et al.
Published: (2025)
Implicit Compressibility of Overparametrized Neural Networks Trained with Heavy-Tailed SGD
by: Wan, Yijun, et al.
Published: (2023)
by: Wan, Yijun, et al.
Published: (2023)
Lower Bounds and Proximally Anchored SGD for Non-Convex Minimization Under Unbounded Variance
by: Fazla, Arda, et al.
Published: (2026)
by: Fazla, Arda, et al.
Published: (2026)
AdaDPIGU: Differentially Private SGD with Adaptive Clipping and Importance-Based Gradient Updates for Deep Neural Networks
by: Zhang, Huiqi, et al.
Published: (2025)
by: Zhang, Huiqi, et al.
Published: (2025)
Optimal Convergence Rates of Deep Neural Network Classifiers
by: Zhang, Zihan, et al.
Published: (2025)
by: Zhang, Zihan, et al.
Published: (2025)
Convergence of SGD for Training Neural Networks with Sliced Wasserstein Losses
by: Tanguy, Eloi
Published: (2023)
by: Tanguy, Eloi
Published: (2023)
Early Directional Convergence in Deep Homogeneous Neural Networks for Small Initializations
by: Kumar, Akshay, et al.
Published: (2024)
by: Kumar, Akshay, et al.
Published: (2024)
How Neural Networks Learn the Support is an Implicit Regularization Effect of SGD
by: Beneventano, Pierfrancesco, et al.
Published: (2024)
by: Beneventano, Pierfrancesco, et al.
Published: (2024)
Neural Variance-aware Dueling Bandits with Deep Representation and Shallow Exploration
by: Oh, Youngmin, et al.
Published: (2025)
by: Oh, Youngmin, et al.
Published: (2025)
Tight Analysis of Decentralized SGD: A Markov Chain Perspective
by: Versini, Lucas, et al.
Published: (2026)
by: Versini, Lucas, et al.
Published: (2026)
Variance-Aware Linear UCB with Deep Representation for Neural Contextual Bandits
by: Bui, Ha Manh, et al.
Published: (2024)
by: Bui, Ha Manh, et al.
Published: (2024)
A Random Matrix Perspective of Echo State Networks: From Precise Bias--Variance Characterization to Optimal Regularization
by: Moakher, Yessin, et al.
Published: (2025)
by: Moakher, Yessin, et al.
Published: (2025)
Shadowheart SGD: Distributed Asynchronous SGD with Optimal Time Complexity Under Arbitrary Computation and Communication Heterogeneity
by: Tyurin, Alexander, et al.
Published: (2024)
by: Tyurin, Alexander, et al.
Published: (2024)
On the Variance of Neural Network Training with respect to Test Sets and Distributions
by: Jordan, Keller
Published: (2023)
by: Jordan, Keller
Published: (2023)
On the Learning Dynamics of Two-layer Linear Networks with Label Noise SGD
by: Zhang, Tongcheng, et al.
Published: (2026)
by: Zhang, Tongcheng, et al.
Published: (2026)
Confidence Interval Construction and Conditional Variance Estimation with Dense ReLU Networks
by: Padilla, Carlos Misael Madrid, et al.
Published: (2024)
by: Padilla, Carlos Misael Madrid, et al.
Published: (2024)
VeLU: Variance-enhanced Learning Unit for Deep Neural Networks
by: Shakarami, Ashkan, et al.
Published: (2025)
by: Shakarami, Ashkan, et al.
Published: (2025)
The Optimality of (Accelerated) SGD for High-Dimensional Quadratic Optimization
by: Zhang, Haihan, et al.
Published: (2024)
by: Zhang, Haihan, et al.
Published: (2024)
Optimal Projection-Free Adaptive SGD for Matrix Optimization
by: Kovalev, Dmitry
Published: (2026)
by: Kovalev, Dmitry
Published: (2026)
Near-Optimal Streaming Heavy-Tailed Statistical Estimation with Clipped SGD
by: Das, Aniket, et al.
Published: (2024)
by: Das, Aniket, et al.
Published: (2024)
Explainable Brain Age Gap Prediction in Neurodegenerative Conditions using coVariance Neural Networks
by: Sihag, Saurabh, et al.
Published: (2025)
by: Sihag, Saurabh, et al.
Published: (2025)
Pruning Deep Convolutional Neural Network Using Conditional Mutual Information
by: Vu-Van, Tien, et al.
Published: (2024)
by: Vu-Van, Tien, et al.
Published: (2024)
Biased Local SGD for Efficient Deep Learning on Heterogeneous Systems
by: Lim, Jihyun, et al.
Published: (2025)
by: Lim, Jihyun, et al.
Published: (2025)
Complexity-Aware Training of Deep Neural Networks for Optimal Structure Discovery
by: Guenter, Valentin Frank Ingmar, et al.
Published: (2024)
by: Guenter, Valentin Frank Ingmar, et al.
Published: (2024)
Suspicious Alignment of SGD: A Fine-Grained Step Size Condition Analysis
by: Deng, Shenyang, et al.
Published: (2026)
by: Deng, Shenyang, et al.
Published: (2026)
Similar Items
-
Weight Initialization and Variance Dynamics in Deep Neural Networks and Large Language Models
by: Han, Yankun
Published: (2025) -
SGD with Partial Hessian for Deep Neural Networks Optimization
by: Sun, Ying, et al.
Published: (2024) -
From SGD to Spectra: A Theory of Neural Network Weight Dynamics
by: Olsen, Brian Richard, et al.
Published: (2025) -
Exploring and Improving Initialization for Deep Graph Neural Networks: A Signal Propagation Perspective
by: Wang, Senmiao, et al.
Published: (2025) -
SGD for Variational Inference: Tackling Unbounded Variance via Preconditioning and Dynamic Batching
by: Labarrière, Hippolyte, et al.
Published: (2026)