Global Convergence and Rich Feature Learning in $L$-Layer Infinite-Width Neural Networks under $μ$P Parametrization
Fuente:
arXiv
Guardado en:
| Autores principales: | Chen, Zixiang, Yang, Greg, Zhao, Qingyue, Gu, Quanquan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Homotopy Relaxation Training Algorithms for Infinite-Width Two-Layer ReLU Neural Networks
por: Yang, Yahong, et al.
Publicado: (2023)
por: Yang, Yahong, et al.
Publicado: (2023)
Matching the Statistical Query Lower Bound for $k$-Sparse Parity Problems with Sign Stochastic Gradient Descent
por: Kou, Yiwen, et al.
Publicado: (2024)
por: Kou, Yiwen, et al.
Publicado: (2024)
A Dual Certificate Approach to Sparsity in Infinite-Width Shallow Neural Networks
por: Del Grande, Leonardo, et al.
Publicado: (2026)
por: Del Grande, Leonardo, et al.
Publicado: (2026)
On the Convergence of Adaptive Gradient Methods for Nonconvex Optimization
por: Zhou, Dongruo, et al.
Publicado: (2018)
por: Zhou, Dongruo, et al.
Publicado: (2018)
Unified Convergence Analysis for Score-Based Diffusion Models with Deterministic Samplers
por: Li, Runjia, et al.
Publicado: (2024)
por: Li, Runjia, et al.
Publicado: (2024)
Over-parameterised Shallow Neural Networks with Asymmetrical Node Scaling: Global Convergence Guarantees and Feature Learning
por: Caron, Francois, et al.
Publicado: (2023)
por: Caron, Francois, et al.
Publicado: (2023)
A Nearly Optimal and Low-Switching Algorithm for Reinforcement Learning with General Function Approximation
por: Zhao, Heyang, et al.
Publicado: (2023)
por: Zhao, Heyang, et al.
Publicado: (2023)
Global Convergence of SGD On Two Layer Neural Nets
por: Gopalani, Pulkit, et al.
Publicado: (2022)
por: Gopalani, Pulkit, et al.
Publicado: (2022)
Dimension-Independent Convergence of Underdamped Langevin Monte Carlo in KL Divergence
por: Zhang, Shiyuan, et al.
Publicado: (2026)
por: Zhang, Shiyuan, et al.
Publicado: (2026)
Grassmannian Geometry and Global Convergence of Variable Projection for Neural Networks
por: Dus, Mathias
Publicado: (2026)
por: Dus, Mathias
Publicado: (2026)
Global Convergence of Four-Layer Matrix Factorization under Random Initialization
por: Luo, Minrui, et al.
Publicado: (2025)
por: Luo, Minrui, et al.
Publicado: (2025)
Global Convergence of SGD For Logistic Loss on Two Layer Neural Nets
por: Gopalani, Pulkit, et al.
Publicado: (2023)
por: Gopalani, Pulkit, et al.
Publicado: (2023)
Pessimistic Nonlinear Least-Squares Value Iteration for Offline Reinforcement Learning
por: Di, Qiwei, et al.
Publicado: (2023)
por: Di, Qiwei, et al.
Publicado: (2023)
Feel-Good Thompson Sampling for Contextual Dueling Bandits
por: Li, Xuheng, et al.
Publicado: (2024)
por: Li, Xuheng, et al.
Publicado: (2024)
Variance-Aware Feel-Good Thompson Sampling for Contextual Bandits
por: Li, Xuheng, et al.
Publicado: (2025)
por: Li, Xuheng, et al.
Publicado: (2025)
Understanding SGD with Exponential Moving Average: A Case Study in Linear Regression
por: Li, Xuheng, et al.
Publicado: (2025)
por: Li, Xuheng, et al.
Publicado: (2025)
Towards a Principled Muon under $μ\mathsf{P}$: Ensuring Spectral Conditions throughout Training
por: Zhao, John
Publicado: (2026)
por: Zhao, John
Publicado: (2026)
Global Contact-Rich Planning with Sparsity-Rich Semidefinite Relaxations
por: Kang, Shucheng, et al.
Publicado: (2025)
por: Kang, Shucheng, et al.
Publicado: (2025)
Nesterov Flow May Travel Infinitely Long to Converge to a Minimizer
por: Ryu, Ernest K.
Publicado: (2026)
por: Ryu, Ernest K.
Publicado: (2026)
Convergence of Implicit Gradient Descent for Training Two-Layer Physics-Informed Neural Networks
por: Xu, Xianliang, et al.
Publicado: (2024)
por: Xu, Xianliang, et al.
Publicado: (2024)
Quantized Distributed Nonconvex Optimization Algorithms with Linear Convergence under the Polyak--$Ł$ojasiewicz Condition
por: Xu, Lei, et al.
Publicado: (2022)
por: Xu, Lei, et al.
Publicado: (2022)
Optimal Horizon-Free Reward-Free Exploration for Linear Mixture MDPs
por: Zhang, Junkai, et al.
Publicado: (2023)
por: Zhang, Junkai, et al.
Publicado: (2023)
MARS-M: When Variance Reduction Meets Matrices
por: Liu, Yifeng, et al.
Publicado: (2025)
por: Liu, Yifeng, et al.
Publicado: (2025)
Convergence of Policy Gradient for Stochastic Linear-Quadratic Control Problem in Infinite Horizon
por: Zhang, Xinpei, et al.
Publicado: (2024)
por: Zhang, Xinpei, et al.
Publicado: (2024)
Convergence Analysis of Two-Layer Neural Networks under Gaussian Input Masking
por: Kolomvaki, Afroditi, et al.
Publicado: (2026)
por: Kolomvaki, Afroditi, et al.
Publicado: (2026)
Policy Gradient Methods for the Cost-Constrained LQR: Strong Duality and Global Convergence
por: Zhao, Feiran, et al.
Publicado: (2024)
por: Zhao, Feiran, et al.
Publicado: (2024)
Non-Parametric Learning of Stochastic Differential Equations with Non-asymptotic Fast Rates of Convergence
por: Bonalli, Riccardo, et al.
Publicado: (2023)
por: Bonalli, Riccardo, et al.
Publicado: (2023)
Convergence Analysis for Learning Orthonormal Deep Linear Neural Networks
por: Qin, Zhen, et al.
Publicado: (2023)
por: Qin, Zhen, et al.
Publicado: (2023)
Physics-Informed Neural Network Policy Iteration: Algorithms, Convergence, and Verification
por: Meng, Yiming, et al.
Publicado: (2024)
por: Meng, Yiming, et al.
Publicado: (2024)
Global Convergence and Error Propagation in Neural Gradient Flows: A Riemannian Optimization Framework
por: Zheng, Shixin, et al.
Publicado: (2026)
por: Zheng, Shixin, et al.
Publicado: (2026)
On the Nonsmooth Geometry and Neural Approximation of the Optimal Value Function of Infinite-Horizon Pendulum Swing-up
por: Han, Haoyu, et al.
Publicado: (2023)
por: Han, Haoyu, et al.
Publicado: (2023)
Optimization for Neural Operators can Benefit from Width
por: Cisneros-Velarde, Pedro, et al.
Publicado: (2025)
por: Cisneros-Velarde, Pedro, et al.
Publicado: (2025)
Shallow Neural Networks Learn Low-Degree Spherical Polynomials with Feature Learning by Learnable Channel Attention
por: Yang, Yingzhen
Publicado: (2025)
por: Yang, Yingzhen
Publicado: (2025)
(Corrected Version) Push-LSVRG-UP: Distributed Stochastic Optimization over Unbalanced Directed Networks with Uncoordinated Triggered Probabilities
por: Hu, Jinhui, et al.
Publicado: (2023)
por: Hu, Jinhui, et al.
Publicado: (2023)
Reinforcement Learning from Human Feedback with Active Queries
por: Ji, Kaixuan, et al.
Publicado: (2024)
por: Ji, Kaixuan, et al.
Publicado: (2024)
Design of Transit Networks: Global Optimization of Continuous Approximation Models via Geometric Programming
por: Mao, Haoyang, et al.
Publicado: (2026)
por: Mao, Haoyang, et al.
Publicado: (2026)
Loss Landscape Characterization of Neural Networks without Over-Parametrization
por: Islamov, Rustem, et al.
Publicado: (2024)
por: Islamov, Rustem, et al.
Publicado: (2024)
Optimality-Informed Neural Networks for Solving Parametric Optimization Problems
por: Hoffmann, Matthias K., et al.
Publicado: (2025)
por: Hoffmann, Matthias K., et al.
Publicado: (2025)
$\ell_{1\text{-}2}$ Regularization for Sparse Optimization: Consistency and Global Convergence
por: Hu, Yaohua, et al.
Publicado: (2026)
por: Hu, Yaohua, et al.
Publicado: (2026)
Learning Parametric Convex Functions
por: Schaller, Maximilian, et al.
Publicado: (2025)
por: Schaller, Maximilian, et al.
Publicado: (2025)
Ejemplares similares
-
Homotopy Relaxation Training Algorithms for Infinite-Width Two-Layer ReLU Neural Networks
por: Yang, Yahong, et al.
Publicado: (2023) -
Matching the Statistical Query Lower Bound for $k$-Sparse Parity Problems with Sign Stochastic Gradient Descent
por: Kou, Yiwen, et al.
Publicado: (2024) -
A Dual Certificate Approach to Sparsity in Infinite-Width Shallow Neural Networks
por: Del Grande, Leonardo, et al.
Publicado: (2026) -
On the Convergence of Adaptive Gradient Methods for Nonconvex Optimization
por: Zhou, Dongruo, et al.
Publicado: (2018) -
Unified Convergence Analysis for Score-Based Diffusion Models with Deterministic Samplers
por: Li, Runjia, et al.
Publicado: (2024)