Weight Initialization and Variance Dynamics in Deep Neural Networks and Large Language Models
Fuente:
arXiv
Salvato in:
| Autore principale: | Han, Yankun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Optimal Condition for Initialization Variance in Deep Neural Networks: An SGD Dynamics Perspective
di: Horii, Hiroshi, et al.
Pubblicazione: (2025)
di: Horii, Hiroshi, et al.
Pubblicazione: (2025)
Optimized Weight Initialization on the Stiefel Manifold for Deep ReLU Neural Networks
di: Lee, Hyungu, et al.
Pubblicazione: (2025)
di: Lee, Hyungu, et al.
Pubblicazione: (2025)
CAWI: Copula-Aligned Weight Initialization for Randomized Neural Networks
di: Akhtar, Mushir, et al.
Pubblicazione: (2026)
di: Akhtar, Mushir, et al.
Pubblicazione: (2026)
Reducing Oversmoothing through Informed Weight Initialization in Graph Neural Networks
di: Kelesis, Dimitrios, et al.
Pubblicazione: (2024)
di: Kelesis, Dimitrios, et al.
Pubblicazione: (2024)
Deep Neural Network Initialization with Sparsity Inducing Activations
di: Price, Ilan, et al.
Pubblicazione: (2024)
di: Price, Ilan, et al.
Pubblicazione: (2024)
Effects of Initialization Biases on Deep Neural Network Training Dynamics
di: Pellegrino, Nicholas, et al.
Pubblicazione: (2025)
di: Pellegrino, Nicholas, et al.
Pubblicazione: (2025)
Posterior Inference on Shallow Infinitely Wide Bayesian Neural Networks under Weights with Unbounded Variance
di: Loría, Jorge, et al.
Pubblicazione: (2023)
di: Loría, Jorge, et al.
Pubblicazione: (2023)
Deep Kernel Posterior Learning under Infinite Variance Prior Weights
di: Loría, Jorge, et al.
Pubblicazione: (2024)
di: Loría, Jorge, et al.
Pubblicazione: (2024)
Robust Weight Initialization for Tanh Neural Networks with Fixed Point Analysis
di: Lee, Hyunwoo, et al.
Pubblicazione: (2024)
di: Lee, Hyunwoo, et al.
Pubblicazione: (2024)
Fair CoVariance Neural Networks
di: Cavallo, Andrea, et al.
Pubblicazione: (2024)
di: Cavallo, Andrea, et al.
Pubblicazione: (2024)
Teasing Apart Architecture and Initial Weights as Sources of Inductive Bias in Neural Networks
di: Bencomo, Gianluca, et al.
Pubblicazione: (2025)
di: Bencomo, Gianluca, et al.
Pubblicazione: (2025)
On the Weight Dynamics of Deep Normalized Networks
di: Mehmeti-Göpel, Christian H. X. Ali, et al.
Pubblicazione: (2023)
di: Mehmeti-Göpel, Christian H. X. Ali, et al.
Pubblicazione: (2023)
LDLT L-Lipschitz Network Weight Parameterization Initialization
di: Juston, Marius F. R., et al.
Pubblicazione: (2026)
di: Juston, Marius F. R., et al.
Pubblicazione: (2026)
DeepWeightFlow: Re-Basined Flow Matching for Generating Neural Network Weights
di: Gupta, Saumya, et al.
Pubblicazione: (2026)
di: Gupta, Saumya, et al.
Pubblicazione: (2026)
Neural Weight Compression for Language Models
di: Ryu, Jegwang, et al.
Pubblicazione: (2025)
di: Ryu, Jegwang, et al.
Pubblicazione: (2025)
Weight-Parameterization in Continuous Time Deep Neural Networks for Surrogate Modeling
di: Rosso, Haley, et al.
Pubblicazione: (2025)
di: Rosso, Haley, et al.
Pubblicazione: (2025)
Supervised Dynamic Dimension Reduction with Deep Neural Network
di: Luo, Zhanye, et al.
Pubblicazione: (2025)
di: Luo, Zhanye, et al.
Pubblicazione: (2025)
Text2Weight: Bridging Natural Language and Neural Network Weight Spaces
di: Tian, Bowen, et al.
Pubblicazione: (2025)
di: Tian, Bowen, et al.
Pubblicazione: (2025)
Graph Neural Network Aided Deep Reinforcement Learning for Resource Allocation in Dynamic Terahertz UAV Networks
di: Hu, Zhifeng, et al.
Pubblicazione: (2025)
di: Hu, Zhifeng, et al.
Pubblicazione: (2025)
Exploring and Improving Initialization for Deep Graph Neural Networks: A Signal Propagation Perspective
di: Wang, Senmiao, et al.
Pubblicazione: (2025)
di: Wang, Senmiao, et al.
Pubblicazione: (2025)
Early Directional Convergence in Deep Homogeneous Neural Networks for Small Initializations
di: Kumar, Akshay, et al.
Pubblicazione: (2024)
di: Kumar, Akshay, et al.
Pubblicazione: (2024)
Neural Variance-aware Dueling Bandits with Deep Representation and Shallow Exploration
di: Oh, Youngmin, et al.
Pubblicazione: (2025)
di: Oh, Youngmin, et al.
Pubblicazione: (2025)
Variance-Aware Linear UCB with Deep Representation for Neural Contextual Bandits
di: Bui, Ha Manh, et al.
Pubblicazione: (2024)
di: Bui, Ha Manh, et al.
Pubblicazione: (2024)
On the Variance of Neural Network Training with respect to Test Sets and Distributions
di: Jordan, Keller
Pubblicazione: (2023)
di: Jordan, Keller
Pubblicazione: (2023)
Learning Guarantee of Reward Modeling Using Deep Neural Networks
di: Luo, Yuanhang, et al.
Pubblicazione: (2025)
di: Luo, Yuanhang, et al.
Pubblicazione: (2025)
VeLU: Variance-enhanced Learning Unit for Deep Neural Networks
di: Shakarami, Ashkan, et al.
Pubblicazione: (2025)
di: Shakarami, Ashkan, et al.
Pubblicazione: (2025)
From SGD to Spectra: A Theory of Neural Network Weight Dynamics
di: Olsen, Brian Richard, et al.
Pubblicazione: (2025)
di: Olsen, Brian Richard, et al.
Pubblicazione: (2025)
Recovering Plasticity of Neural Networks via Soft Weight Rescaling
di: Oh, Seungwon, et al.
Pubblicazione: (2025)
di: Oh, Seungwon, et al.
Pubblicazione: (2025)
Geometric Flow Models over Neural Network Weights
di: Erdogan, Ege
Pubblicazione: (2025)
di: Erdogan, Ege
Pubblicazione: (2025)
WAVE: Weight Templates for Adaptive Initialization of Variable-sized Models
di: Feng, Fu, et al.
Pubblicazione: (2024)
di: Feng, Fu, et al.
Pubblicazione: (2024)
VCRL: Variance-based Curriculum Reinforcement Learning for Large Language Models
di: Jiang, Guochao, et al.
Pubblicazione: (2025)
di: Jiang, Guochao, et al.
Pubblicazione: (2025)
Dynamic Weight Adjusting Deep Q-Networks for Real-Time Environmental Adaptation
di: Zhang, Xinhao, et al.
Pubblicazione: (2024)
di: Zhang, Xinhao, et al.
Pubblicazione: (2024)
Variance-Aware Adaptive Weighting for Diffusion Model Training
di: Sun, Nanlong, et al.
Pubblicazione: (2026)
di: Sun, Nanlong, et al.
Pubblicazione: (2026)
Wormhole Dynamics in Deep Neural Networks
di: Lai, Yen-Lung, et al.
Pubblicazione: (2025)
di: Lai, Yen-Lung, et al.
Pubblicazione: (2025)
Principal Components for Neural Network Initialization
di: Phan, Nhan, et al.
Pubblicazione: (2025)
di: Phan, Nhan, et al.
Pubblicazione: (2025)
The SkipSponge Attack: Sponge Weight Poisoning of Deep Neural Networks
di: Lintelo, Jona te, et al.
Pubblicazione: (2024)
di: Lintelo, Jona te, et al.
Pubblicazione: (2024)
Deep Hierarchical Learning with Nested Subspace Networks for Large Language Models
di: Rauba, Paulius, et al.
Pubblicazione: (2025)
di: Rauba, Paulius, et al.
Pubblicazione: (2025)
LLaDA 1.5: Variance-Reduced Preference Optimization for Large Language Diffusion Models
di: Zhu, Fengqi, et al.
Pubblicazione: (2025)
di: Zhu, Fengqi, et al.
Pubblicazione: (2025)
Variance-aware Reward Modeling with Anchor Guidance
di: Fang, Shuxing, et al.
Pubblicazione: (2026)
di: Fang, Shuxing, et al.
Pubblicazione: (2026)
Cooperative Variance Estimation and Bayesian Neural Networks for Disentangling Aleatoric and Epistemic Uncertainties
di: Yi, Jiaxiang, et al.
Pubblicazione: (2025)
di: Yi, Jiaxiang, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Optimal Condition for Initialization Variance in Deep Neural Networks: An SGD Dynamics Perspective
di: Horii, Hiroshi, et al.
Pubblicazione: (2025) -
Optimized Weight Initialization on the Stiefel Manifold for Deep ReLU Neural Networks
di: Lee, Hyungu, et al.
Pubblicazione: (2025) -
CAWI: Copula-Aligned Weight Initialization for Randomized Neural Networks
di: Akhtar, Mushir, et al.
Pubblicazione: (2026) -
Reducing Oversmoothing through Informed Weight Initialization in Graph Neural Networks
di: Kelesis, Dimitrios, et al.
Pubblicazione: (2024) -
Deep Neural Network Initialization with Sparsity Inducing Activations
di: Price, Ilan, et al.
Pubblicazione: (2024)