Precise gradient descent training dynamics for finite-width multi-layer neural networks
Fuente:
arXiv
Salvato in:
| Autori principali: | Han, Qiyang, Imaizumi, Masaaki |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Long-time dynamics and universality of nonconvex gradient descent
di: Han, Qiyang
Pubblicazione: (2025)
di: Han, Qiyang
Pubblicazione: (2025)
Gradient descent inference in empirical risk minimization
di: Han, Qiyang, et al.
Pubblicazione: (2024)
di: Han, Qiyang, et al.
Pubblicazione: (2024)
Convergence of flow-based generative models via proximal gradient descent in Wasserstein space
di: Cheng, Xiuyuan, et al.
Pubblicazione: (2023)
di: Cheng, Xiuyuan, et al.
Pubblicazione: (2023)
Optimism Stabilizes Thompson Sampling for Adaptive Inference
di: Yan, Shunxing, et al.
Pubblicazione: (2026)
di: Yan, Shunxing, et al.
Pubblicazione: (2026)
Inverse Mixed-Integer Programming: Learning Constraints then Objective Functions
di: Kitaoka, Akira
Pubblicazione: (2025)
di: Kitaoka, Akira
Pubblicazione: (2025)
Learning to Fuse Temporal Proximity Networks: A Case Study in Chimpanzee Social Interactions
di: He, Yixuan, et al.
Pubblicazione: (2025)
di: He, Yixuan, et al.
Pubblicazione: (2025)
FraPPE: Fast and Efficient Preference-based Pure Exploration
di: Das, Udvas, et al.
Pubblicazione: (2025)
di: Das, Udvas, et al.
Pubblicazione: (2025)
Statistical and Algorithmic Foundations of Reinforcement Learning
di: Chi, Yuejie, et al.
Pubblicazione: (2025)
di: Chi, Yuejie, et al.
Pubblicazione: (2025)
Decoupled Continuous-Time Reinforcement Learning via Hamiltonian Flow
di: Nguyen, Minh
Pubblicazione: (2026)
di: Nguyen, Minh
Pubblicazione: (2026)
Byzantine Machine Learning: MultiKrum and an optimal notion of robustness
di: Bareilles, Gilles, et al.
Pubblicazione: (2026)
di: Bareilles, Gilles, et al.
Pubblicazione: (2026)
Sinkhorn Based Associative Memory Retrieval Using Spherical Hellinger Kantorovich Dynamics
di: Mustafi, Aratrika, et al.
Pubblicazione: (2026)
di: Mustafi, Aratrika, et al.
Pubblicazione: (2026)
Straight-Through meets Sparse Recovery: the Support Exploration Algorithm
di: Mohamed, Mimoun, et al.
Pubblicazione: (2023)
di: Mohamed, Mimoun, et al.
Pubblicazione: (2023)
A Differential and Pointwise Control Approach to Reinforcement Learning
di: Nguyen, Minh, et al.
Pubblicazione: (2024)
di: Nguyen, Minh, et al.
Pubblicazione: (2024)
Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality
di: Chen, Siyu, et al.
Pubblicazione: (2024)
di: Chen, Siyu, et al.
Pubblicazione: (2024)
Sail into the Headwind: Alignment via Robust Rewards and Dynamic Labels against Reward Hacking
di: Rashidinejad, Paria, et al.
Pubblicazione: (2024)
di: Rashidinejad, Paria, et al.
Pubblicazione: (2024)
Smooth Non-Stationary Bandits
di: Jia, Su, et al.
Pubblicazione: (2023)
di: Jia, Su, et al.
Pubblicazione: (2023)
Piecewise Polynomial Regression of Tame Functions via Integer Programming
di: Bareilles, Gilles, et al.
Pubblicazione: (2023)
di: Bareilles, Gilles, et al.
Pubblicazione: (2023)
Fast Spawn\&Prune (FS\&P): Global convergence of stochastic conic particle gradient descent via birth/death process
di: De Castro, Yohann, et al.
Pubblicazione: (2026)
di: De Castro, Yohann, et al.
Pubblicazione: (2026)
Recurrent neural networks: vanishing and exploding gradients are not the end of the story
di: Zucchet, Nicolas, et al.
Pubblicazione: (2024)
di: Zucchet, Nicolas, et al.
Pubblicazione: (2024)
When majority rules, minority loses: bias amplification of gradient descent
di: Bachoc, François, et al.
Pubblicazione: (2025)
di: Bachoc, François, et al.
Pubblicazione: (2025)
Pinet: Optimizing hard-constrained neural networks with orthogonal projection layers
di: Grontas, Panagiotis D., et al.
Pubblicazione: (2025)
di: Grontas, Panagiotis D., et al.
Pubblicazione: (2025)
Joint learning of a network of linear dynamical systems via total variation penalization
di: Donnat, Claire, et al.
Pubblicazione: (2025)
di: Donnat, Claire, et al.
Pubblicazione: (2025)
Efficient and provably convergent end-to-end training of deep neural networks with linear constraints
di: Yang, Zonglin, et al.
Pubblicazione: (2026)
di: Yang, Zonglin, et al.
Pubblicazione: (2026)
The duality structure gradient descent algorithm: analysis and applications to neural networks
di: Flynn, Thomas
Pubblicazione: (2017)
di: Flynn, Thomas
Pubblicazione: (2017)
Robust stochastic first order methods in heavy-tailed noise via medoid mini-batch gradient sampling
di: Vukovic, Manojlo, et al.
Pubblicazione: (2026)
di: Vukovic, Manojlo, et al.
Pubblicazione: (2026)
State evolution beyond first-order methods I: Rigorous predictions and finite-sample guarantees
di: Celentano, Michael, et al.
Pubblicazione: (2025)
di: Celentano, Michael, et al.
Pubblicazione: (2025)
Convergence of continuous-time stochastic gradient descent with applications to deep neural networks
di: Lugosi, Gabor, et al.
Pubblicazione: (2024)
di: Lugosi, Gabor, et al.
Pubblicazione: (2024)
Optimal transport natural gradient for statistical manifolds with continuous sample space
di: Chen, Yifan, et al.
Pubblicazione: (2018)
di: Chen, Yifan, et al.
Pubblicazione: (2018)
Non-convergence to global minimizers for Adam and stochastic gradient descent optimization and constructions of local minimizers in the training of artificial neural networks
di: Jentzen, Arnulf, et al.
Pubblicazione: (2024)
di: Jentzen, Arnulf, et al.
Pubblicazione: (2024)
Geometry-induced Regularization in Deep ReLU Neural Networks
di: Bona-Pellissier, Joachim, et al.
Pubblicazione: (2024)
di: Bona-Pellissier, Joachim, et al.
Pubblicazione: (2024)
Learning an Optimal Assortment Policy under Observational Data
di: Han, Yuxuan, et al.
Pubblicazione: (2025)
di: Han, Yuxuan, et al.
Pubblicazione: (2025)
Robust Assortment Optimization from Observational Data
di: Lu, Miao, et al.
Pubblicazione: (2026)
di: Lu, Miao, et al.
Pubblicazione: (2026)
Learning linear dynamical systems under convex constraints
di: Tyagi, Hemant, et al.
Pubblicazione: (2023)
di: Tyagi, Hemant, et al.
Pubblicazione: (2023)
Convergence of stochastic gradient descent under a local Lojasiewicz condition for deep neural networks
di: An, Jing, et al.
Pubblicazione: (2023)
di: An, Jing, et al.
Pubblicazione: (2023)
The late-stage training dynamics of (stochastic) subgradient descent on homogeneous neural networks
di: Schechtman, Sholom, et al.
Pubblicazione: (2025)
di: Schechtman, Sholom, et al.
Pubblicazione: (2025)
CITE: Anytime-Valid Statistical Inference in LLM Self-Consistency
di: Ota, Hirofumi, et al.
Pubblicazione: (2026)
di: Ota, Hirofumi, et al.
Pubblicazione: (2026)
Thompson sampling: Precise arm-pull dynamics and adaptive inference
di: Han, Qiyang
Pubblicazione: (2026)
di: Han, Qiyang
Pubblicazione: (2026)
Mean-field underdamped Langevin dynamics and its spacetime discretization
di: Fu, Qiang, et al.
Pubblicazione: (2023)
di: Fu, Qiang, et al.
Pubblicazione: (2023)
A multiobjective continuation method to compute the regularization path of deep neural networks
di: Amakor, Augustina C., et al.
Pubblicazione: (2023)
di: Amakor, Augustina C., et al.
Pubblicazione: (2023)
Conformal Prediction in The Loop: A Feedback-Based Uncertainty Model for Trajectory Optimization
di: Wang, Han, et al.
Pubblicazione: (2025)
di: Wang, Han, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Long-time dynamics and universality of nonconvex gradient descent
di: Han, Qiyang
Pubblicazione: (2025) -
Gradient descent inference in empirical risk minimization
di: Han, Qiyang, et al.
Pubblicazione: (2024) -
Convergence of flow-based generative models via proximal gradient descent in Wasserstein space
di: Cheng, Xiuyuan, et al.
Pubblicazione: (2023) -
Optimism Stabilizes Thompson Sampling for Adaptive Inference
di: Yan, Shunxing, et al.
Pubblicazione: (2026) -
Inverse Mixed-Integer Programming: Learning Constraints then Objective Functions
di: Kitaoka, Akira
Pubblicazione: (2025)