Noise Balance and Stationary Distribution of Stochastic Gradient Descent
Fuente:
arXiv
Guardado en:
| Autores principales: | Ziyin, Liu, Li, Hongchao, Ueda, Masahito |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Parameter Symmetry and Noise Equilibrium of Stochastic Gradient Descent
por: Ziyin, Liu, et al.
Publicado: (2024)
por: Ziyin, Liu, et al.
Publicado: (2024)
Type-II Saddles and Probabilistic Stability of Stochastic Gradient Descent
por: Ziyin, Liu, et al.
Publicado: (2023)
por: Ziyin, Liu, et al.
Publicado: (2023)
Stochastic Re-weighted Gradient Descent via Distributionally Robust Optimization
por: Kumar, Ramnath, et al.
Publicado: (2023)
por: Kumar, Ramnath, et al.
Publicado: (2023)
Trustworthiness of Stochastic Gradient Descent in Distributed Learning
por: Li, Hongyang, et al.
Publicado: (2024)
por: Li, Hongyang, et al.
Publicado: (2024)
Adaptive Batch Sizes Using Non-Euclidean Gradient Noise Scales for Stochastic Sign and Spectral Descent
por: Naganuma, Hiroki, et al.
Publicado: (2026)
por: Naganuma, Hiroki, et al.
Publicado: (2026)
Adaptive Heavy-Tailed Stochastic Gradient Descent
por: Gong, Bodu, et al.
Publicado: (2025)
por: Gong, Bodu, et al.
Publicado: (2025)
Stochastic Gradient Descent with Momentum is Algorithmically Stable
por: Lei, Yunwen, et al.
Publicado: (2026)
por: Lei, Yunwen, et al.
Publicado: (2026)
Can LLMs predict the convergence of Stochastic Gradient Descent?
por: Zekri, Oussama, et al.
Publicado: (2024)
por: Zekri, Oussama, et al.
Publicado: (2024)
PSMGD: Periodic Stochastic Multi-Gradient Descent for Fast Multi-Objective Optimization
por: Xu, Mingjing, et al.
Publicado: (2024)
por: Xu, Mingjing, et al.
Publicado: (2024)
Almost Bayesian: The Fractal Dynamics of Stochastic Gradient Descent
por: Hennick, Max, et al.
Publicado: (2025)
por: Hennick, Max, et al.
Publicado: (2025)
On the Convergence of (Stochastic) Gradient Descent for Kolmogorov--Arnold Networks
por: Gao, Yihang, et al.
Publicado: (2024)
por: Gao, Yihang, et al.
Publicado: (2024)
Unveiling m-Sharpness Through the Structure of Stochastic Gradient Noise
por: Luo, Haocheng, et al.
Publicado: (2025)
por: Luo, Haocheng, et al.
Publicado: (2025)
Randomness and Interpolation Improve Gradient Descent
por: Li, Jiawen, et al.
Publicado: (2025)
por: Li, Jiawen, et al.
Publicado: (2025)
Conflict-Averse Gradient Descent for Multi-task Learning
por: Liu, Bo, et al.
Publicado: (2021)
por: Liu, Bo, et al.
Publicado: (2021)
Gradient Descent Algorithm Survey
por: Fucheng, Deng, et al.
Publicado: (2025)
por: Fucheng, Deng, et al.
Publicado: (2025)
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization
por: Abuduweili, Abulikemu, et al.
Publicado: (2024)
por: Abuduweili, Abulikemu, et al.
Publicado: (2024)
Weighted Low-rank Approximation via Stochastic Gradient Descent on Manifolds
por: Xu, Conglong, et al.
Publicado: (2025)
por: Xu, Conglong, et al.
Publicado: (2025)
Incentivized Exploration of Non-Stationary Stochastic Bandits
por: Chakraborty, Sourav, et al.
Publicado: (2024)
por: Chakraborty, Sourav, et al.
Publicado: (2024)
Learning Associative Memories with Gradient Descent
por: Cabannes, Vivien, et al.
Publicado: (2024)
por: Cabannes, Vivien, et al.
Publicado: (2024)
ONG: Orthogonal Natural Gradient Descent
por: Yadav, Yajat, et al.
Publicado: (2025)
por: Yadav, Yajat, et al.
Publicado: (2025)
Reconstructing Deep Neural Networks: Unleashing the Optimization Potential of Natural Gradient Descent
por: Liu, Weihua, et al.
Publicado: (2024)
por: Liu, Weihua, et al.
Publicado: (2024)
Three Mechanisms of Feature Learning in a Linear Network
por: Xu, Yizhou, et al.
Publicado: (2024)
por: Xu, Yizhou, et al.
Publicado: (2024)
FedBCD:Communication-Efficient Accelerated Block Coordinate Gradient Descent for Federated Learning
por: Liu, Junkang, et al.
Publicado: (2026)
por: Liu, Junkang, et al.
Publicado: (2026)
Elastic Multi-Gradient Descent for Parallel Continual Learning
por: Lyu, Fan, et al.
Publicado: (2024)
por: Lyu, Fan, et al.
Publicado: (2024)
Stochastic Collapse: How Gradient Noise Attracts SGD Dynamics Towards Simpler Subnetworks
por: Chen, Feng, et al.
Publicado: (2023)
por: Chen, Feng, et al.
Publicado: (2023)
Partition Tree Weighting for Non-Stationary Stochastic Bandits
por: Veness, Joel, et al.
Publicado: (2025)
por: Veness, Joel, et al.
Publicado: (2025)
Vanilla Gradient Descent for Oblique Decision Trees
por: Panda, Subrat Prasad, et al.
Publicado: (2024)
por: Panda, Subrat Prasad, et al.
Publicado: (2024)
Revealing Modular Gradient Noise Imbalance in LLMs: Calibrating Adam via Signal-to-Noise Ratio
por: Wen, Ziqing, et al.
Publicado: (2026)
por: Wen, Ziqing, et al.
Publicado: (2026)
Efficient Search for Customized Activation Functions with Gradient Descent
por: Strack, Lukas, et al.
Publicado: (2024)
por: Strack, Lukas, et al.
Publicado: (2024)
The Initialization Determines Whether In-Context Learning Is Gradient Descent
por: Xie, Shifeng, et al.
Publicado: (2025)
por: Xie, Shifeng, et al.
Publicado: (2025)
Enhancing Stochastic Gradient Descent: A Unified Framework and Novel Acceleration Methods for Faster Convergence
por: Deng, Yichuan, et al.
Publicado: (2024)
por: Deng, Yichuan, et al.
Publicado: (2024)
Gradient Descent Efficiency Index
por: Dhingra, Aviral
Publicado: (2024)
por: Dhingra, Aviral
Publicado: (2024)
Fisher-Orthogonal Projected Natural Gradient Descent for Continual Learning
por: Garg, Ishir, et al.
Publicado: (2026)
por: Garg, Ishir, et al.
Publicado: (2026)
A Universal Banach--Bregman Framework for Stochastic Iterations: Unifying Stochastic Mirror Descent, Learning and LLM Training
por: Zhang, Johnny R., et al.
Publicado: (2025)
por: Zhang, Johnny R., et al.
Publicado: (2025)
Clarifying Shampoo: Adapting Spectral Descent to Stochasticity and the Parameter Trajectory
por: Eschenhagen, Runa, et al.
Publicado: (2026)
por: Eschenhagen, Runa, et al.
Publicado: (2026)
Why Does Stochastic Gradient Descent Slow Down in Low-Precision Training?
por: Yun, Vincent-Daniel
Publicado: (2025)
por: Yun, Vincent-Daniel
Publicado: (2025)
Remove Symmetries to Control Model Expressivity and Improve Optimization
por: Ziyin, Liu, et al.
Publicado: (2024)
por: Ziyin, Liu, et al.
Publicado: (2024)
Stochastic Resetting Mitigates Latent Gradient Bias of SGD from Label Noise
por: Bae, Youngkyoung, et al.
Publicado: (2024)
por: Bae, Youngkyoung, et al.
Publicado: (2024)
GradTree: Learning Axis-Aligned Decision Trees with Gradient Descent
por: Marton, Sascha, et al.
Publicado: (2023)
por: Marton, Sascha, et al.
Publicado: (2023)
Geometrically Inspired Kernel Machines for Collaborative Learning Beyond Gradient Descent
por: Kumar, Mohit, et al.
Publicado: (2024)
por: Kumar, Mohit, et al.
Publicado: (2024)
Ejemplares similares
-
Parameter Symmetry and Noise Equilibrium of Stochastic Gradient Descent
por: Ziyin, Liu, et al.
Publicado: (2024) -
Type-II Saddles and Probabilistic Stability of Stochastic Gradient Descent
por: Ziyin, Liu, et al.
Publicado: (2023) -
Stochastic Re-weighted Gradient Descent via Distributionally Robust Optimization
por: Kumar, Ramnath, et al.
Publicado: (2023) -
Trustworthiness of Stochastic Gradient Descent in Distributed Learning
por: Li, Hongyang, et al.
Publicado: (2024) -
Adaptive Batch Sizes Using Non-Euclidean Gradient Noise Scales for Stochastic Sign and Spectral Descent
por: Naganuma, Hiroki, et al.
Publicado: (2026)