Beyond Gaussian Initializations: Signal Preserving Weight Initialization for Odd-Sigmoid Activations
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Hyunwoo, Choi, Hayoung, Kim, Hyunju |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Robust Weight Initialization for Tanh Neural Networks with Fixed Point Analysis
by: Lee, Hyunwoo, et al.
Published: (2024)
by: Lee, Hyunwoo, et al.
Published: (2024)
Progressive Weight Loading: Accelerating Initial Inference and Gradually Boosting Performance on Resource-Constrained Environments
by: Kim, Hyunwoo, et al.
Published: (2025)
by: Kim, Hyunwoo, et al.
Published: (2025)
Optimized Weight Initialization on the Stiefel Manifold for Deep ReLU Neural Networks
by: Lee, Hyungu, et al.
Published: (2025)
by: Lee, Hyungu, et al.
Published: (2025)
PINNet: a deep neural network with pathway prior knowledge for Alzheimer's disease
by: Kim, Yeojin, et al.
Published: (2022)
by: Kim, Yeojin, et al.
Published: (2022)
Deep Metric Loss for Multimodal Learning
by: Moon, Sehwan, et al.
Published: (2023)
by: Moon, Sehwan, et al.
Published: (2023)
Teasing Apart Architecture and Initial Weights as Sources of Inductive Bias in Neural Networks
by: Bencomo, Gianluca, et al.
Published: (2025)
by: Bencomo, Gianluca, et al.
Published: (2025)
Fast Training of Sinusoidal Neural Fields via Scaling Initialization
by: Yeom, Taesun, et al.
Published: (2024)
by: Yeom, Taesun, et al.
Published: (2024)
Frictional Q-Learning
by: Kim, Hyunwoo, et al.
Published: (2025)
by: Kim, Hyunwoo, et al.
Published: (2025)
Entropy-Aware Model Initialization for Effective Exploration in Deep Reinforcement Learning
by: Jang, Sooyoung, et al.
Published: (2021)
by: Jang, Sooyoung, et al.
Published: (2021)
Spatio-Temporal Graphs Beyond Grids: Benchmark for Maritime Anomaly Detection
by: Kim, Jeehong, et al.
Published: (2025)
by: Kim, Jeehong, et al.
Published: (2025)
FlowerFormer: Empowering Neural Architecture Encoding using a Flow-aware Graph Transformer
by: Hwang, Dongyeong, et al.
Published: (2024)
by: Hwang, Dongyeong, et al.
Published: (2024)
Principal Components for Neural Network Initialization
by: Phan, Nhan, et al.
Published: (2025)
by: Phan, Nhan, et al.
Published: (2025)
LION-DG: Layer-Informed Initialization with Deep Gradient Protocols for Accelerated Neural Network Training
by: Kim, Hyunjun
Published: (2026)
by: Kim, Hyunjun
Published: (2026)
Preserving Bilinear Weight Spectra with a Signed and Shrunk Quadratic Activation Function
by: Abohwo, Jason, et al.
Published: (2025)
by: Abohwo, Jason, et al.
Published: (2025)
Revisiting Weight Averaging for Model Merging
by: Choi, Jiho, et al.
Published: (2024)
by: Choi, Jiho, et al.
Published: (2024)
Global Minimizers of Sigmoid Contrastive Loss
by: Bangachev, Kiril, et al.
Published: (2025)
by: Bangachev, Kiril, et al.
Published: (2025)
Geometric Regularization in Mixture-of-Experts: The Disconnect Between Weights and Activations
by: Kim, Hyunjun
Published: (2026)
by: Kim, Hyunjun
Published: (2026)
Achieving the Tightest Relaxation of Sigmoids for Formal Verification
by: Chevalier, Samuel, et al.
Published: (2024)
by: Chevalier, Samuel, et al.
Published: (2024)
Revisiting Glorot Initialization for Long-Range Linear Recurrences
by: Bar, Noga, et al.
Published: (2025)
by: Bar, Noga, et al.
Published: (2025)
When Bias Meets Trainability: Connecting Theories of Initialization
by: Bassi, Alberto, et al.
Published: (2025)
by: Bassi, Alberto, et al.
Published: (2025)
The Initialization Determines Whether In-Context Learning Is Gradient Descent
by: Xie, Shifeng, et al.
Published: (2025)
by: Xie, Shifeng, et al.
Published: (2025)
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization
by: Abuduweili, Abulikemu, et al.
Published: (2024)
by: Abuduweili, Abulikemu, et al.
Published: (2024)
Initializing Services in Interactive ML Systems for Diverse Users
by: Bose, Avinandan, et al.
Published: (2023)
by: Bose, Avinandan, et al.
Published: (2023)
NNiT: Width-Agnostic Neural Network Generation with Structurally Aligned Weight Spaces
by: Kim, Jiwoo, et al.
Published: (2026)
by: Kim, Jiwoo, et al.
Published: (2026)
World Models for Autonomous Driving: An Initial Survey
by: Guan, Yanchen, et al.
Published: (2024)
by: Guan, Yanchen, et al.
Published: (2024)
Bridging KAN and MLP: MJKAN, a Hybrid Architecture with Both Efficiency and Expressiveness
by: Joo, Hanseon, et al.
Published: (2025)
by: Joo, Hanseon, et al.
Published: (2025)
UNR-Explainer: Counterfactual Explanations for Unsupervised Node Representation Learning Models
by: Kang, Hyunju, et al.
Published: (2026)
by: Kang, Hyunju, et al.
Published: (2026)
Random Initialization of Gated Sparse Adapters
by: Retault, Vi, et al.
Published: (2025)
by: Retault, Vi, et al.
Published: (2025)
NoiseAR: AutoRegressing Initial Noise Prior for Diffusion Models
by: Li, Zeming, et al.
Published: (2025)
by: Li, Zeming, et al.
Published: (2025)
IDInit: A Universal and Stable Initialization Method for Neural Network Training
by: Pan, Yu, et al.
Published: (2025)
by: Pan, Yu, et al.
Published: (2025)
Randomly Initialized Networks Can Learn from Peer-to-Peer Consensus
by: Rodríguez-Betancourt, Esteban, et al.
Published: (2026)
by: Rodríguez-Betancourt, Esteban, et al.
Published: (2026)
Bit-Identical Medical Deep Learning via Structured Orthogonal Initialization
by: Shkolnikov, Yakov Pyotr
Published: (2026)
by: Shkolnikov, Yakov Pyotr
Published: (2026)
Rapid Deployment of DNNs for Edge Computing via Structured Pruning at Initialization
by: Eccles, Bailey J., et al.
Published: (2024)
by: Eccles, Bailey J., et al.
Published: (2024)
No More Adam: Learning Rate Scaling at Initialization is All You Need
by: Xu, Minghao, et al.
Published: (2024)
by: Xu, Minghao, et al.
Published: (2024)
Inversion-based Latent Bayesian Optimization
by: Chu, Jaewon, et al.
Published: (2024)
by: Chu, Jaewon, et al.
Published: (2024)
Mimetic Initialization of MLPs
by: Trockman, Asher, et al.
Published: (2026)
by: Trockman, Asher, et al.
Published: (2026)
JUDO: A Juxtaposed Domain-Oriented Multimodal Reasoner for Industrial Anomaly QA
by: Kang, Hyunju, et al.
Published: (2026)
by: Kang, Hyunju, et al.
Published: (2026)
Exploring Learngene via Stage-wise Weight Sharing for Initializing Variable-sized Models
by: Xia, Shi-Yu, et al.
Published: (2024)
by: Xia, Shi-Yu, et al.
Published: (2024)
Sparse Weight Averaging with Multiple Particles for Iterative Magnitude Pruning
by: Choi, Moonseok, et al.
Published: (2023)
by: Choi, Moonseok, et al.
Published: (2023)
The Impact of Initialization on LoRA Finetuning Dynamics
by: Hayou, Soufiane, et al.
Published: (2024)
by: Hayou, Soufiane, et al.
Published: (2024)
Similar Items
-
Robust Weight Initialization for Tanh Neural Networks with Fixed Point Analysis
by: Lee, Hyunwoo, et al.
Published: (2024) -
Progressive Weight Loading: Accelerating Initial Inference and Gradually Boosting Performance on Resource-Constrained Environments
by: Kim, Hyunwoo, et al.
Published: (2025) -
Optimized Weight Initialization on the Stiefel Manifold for Deep ReLU Neural Networks
by: Lee, Hyungu, et al.
Published: (2025) -
PINNet: a deep neural network with pathway prior knowledge for Alzheimer's disease
by: Kim, Yeojin, et al.
Published: (2022) -
Deep Metric Loss for Multimodal Learning
by: Moon, Sehwan, et al.
Published: (2023)