Convergence Analysis for Learning Orthonormal Deep Linear Neural Networks
Fuente:
arXiv
Saved in:
| Main Authors: | Qin, Zhen, Tan, Xuwei, Zhu, Zhihui |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the Convergence of Gradient Descent on Learning Transformers with Residual Connections
by: Qin, Zhen, et al.
Published: (2025)
by: Qin, Zhen, et al.
Published: (2025)
Learning to Adapt: In-Context Learning Beyond Stationarity
by: Qin, Zhen, et al.
Published: (2026)
by: Qin, Zhen, et al.
Published: (2026)
Robust Low-rank Tensor Train Recovery
by: Qin, Zhen, et al.
Published: (2024)
by: Qin, Zhen, et al.
Published: (2024)
Computational and Statistical Guarantees for Tensor-on-Tensor Regression with Tensor Train Decomposition
by: Qin, Zhen, et al.
Published: (2024)
by: Qin, Zhen, et al.
Published: (2024)
Data Collaboration Analysis with Orthonormal Basis Selection and Alignment
by: Nosaka, Keiyu, et al.
Published: (2024)
by: Nosaka, Keiyu, et al.
Published: (2024)
Solving Inverse Problems with Deep Linear Neural Networks: Global Convergence Guarantees for Gradient Descent with Weight Decay
by: Laus, Hannah, et al.
Published: (2025)
by: Laus, Hannah, et al.
Published: (2025)
Neural Collapse under Gradient Flow on Shallow ReLU Networks for Orthogonally Separable Data
by: Min, Hancheng, et al.
Published: (2025)
by: Min, Hancheng, et al.
Published: (2025)
Linear-Quadratic Mean-Field Reinforcement Learning: Convergence of Policy Gradient Methods
by: Carmona, René, et al.
Published: (2019)
by: Carmona, René, et al.
Published: (2019)
Early Directional Convergence in Deep Homogeneous Neural Networks for Small Initializations
by: Kumar, Akshay, et al.
Published: (2024)
by: Kumar, Akshay, et al.
Published: (2024)
Provable Accelerated Convergence of Nesterov's Momentum for Deep ReLU Neural Networks
by: Liao, Fangshuo, et al.
Published: (2023)
by: Liao, Fangshuo, et al.
Published: (2023)
Guaranteed Nonconvex Factorization Approach for Tensor Train Recovery
by: Qin, Zhen, et al.
Published: (2024)
by: Qin, Zhen, et al.
Published: (2024)
A Scalable Factorization Approach for High-Order Structured Tensor Recovery
by: Qin, Zhen, et al.
Published: (2025)
by: Qin, Zhen, et al.
Published: (2025)
Dion: Distributed Orthonormalized Updates
by: Ahn, Kwangjun, et al.
Published: (2025)
by: Ahn, Kwangjun, et al.
Published: (2025)
Convergence of Gradient Descent for Recurrent Neural Networks: A Nonasymptotic Analysis
by: Cayci, Semih, et al.
Published: (2024)
by: Cayci, Semih, et al.
Published: (2024)
Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks
by: Beneventano, Pierfrancesco, et al.
Published: (2025)
by: Beneventano, Pierfrancesco, et al.
Published: (2025)
Communication Efficient Federated Learning with Linear Convergence on Heterogeneous Data
by: Liu, Jie, et al.
Published: (2025)
by: Liu, Jie, et al.
Published: (2025)
An Adaptive and Stability-Promoting Layerwise Training Approach for Sparse Deep Neural Network Architecture
by: Krishnanunni, C G, et al.
Published: (2022)
by: Krishnanunni, C G, et al.
Published: (2022)
Optimization Insights into Deep Diagonal Linear Networks
by: Labarrière, Hippolyte, et al.
Published: (2024)
by: Labarrière, Hippolyte, et al.
Published: (2024)
Over-parameterised Shallow Neural Networks with Asymmetrical Node Scaling: Global Convergence Guarantees and Feature Learning
by: Caron, Francois, et al.
Published: (2023)
by: Caron, Francois, et al.
Published: (2023)
BOOOM: Loss-Function-Agnostic Black-Box Optimization over Orthonormal Manifolds for Machine Learning and Statistical Inference
by: Kim, Beomchang, et al.
Published: (2026)
by: Kim, Beomchang, et al.
Published: (2026)
Convergence Analysis of the Wasserstein Proximal Algorithm beyond Geodesic Convexity
by: Zhu, Shuailong, et al.
Published: (2025)
by: Zhu, Shuailong, et al.
Published: (2025)
Nearly Optimal Linear Convergence of Stochastic Primal-Dual Methods for Linear Programming
by: Lu, Haihao, et al.
Published: (2021)
by: Lu, Haihao, et al.
Published: (2021)
Linear Convergence of Entropy-Regularized Natural Policy Gradient with Linear Function Approximation
by: Cayci, Semih, et al.
Published: (2021)
by: Cayci, Semih, et al.
Published: (2021)
Directional Convergence Near Small Initializations and Saddles in Two-Homogeneous Neural Networks
by: Kumar, Akshay, et al.
Published: (2024)
by: Kumar, Akshay, et al.
Published: (2024)
QLABGrad: a Hyperparameter-Free and Convergence-Guaranteed Scheme for Deep Learning
by: Fu, Minghan, et al.
Published: (2023)
by: Fu, Minghan, et al.
Published: (2023)
Active Learning of Deep Neural Networks via Gradient-Free Cutting Planes
by: Zhang, Erica, et al.
Published: (2024)
by: Zhang, Erica, et al.
Published: (2024)
CRONOS: Enhancing Deep Learning with Scalable GPU Accelerated Convex Neural Networks
by: Feng, Miria, et al.
Published: (2024)
by: Feng, Miria, et al.
Published: (2024)
Local Linear Convergence of Infeasible Optimization with Orthogonal Constraints
by: Sun, Youbang, et al.
Published: (2024)
by: Sun, Youbang, et al.
Published: (2024)
Physics-Informed Neural Networks with Hard Linear Equality Constraints
by: Chen, Hao, et al.
Published: (2024)
by: Chen, Hao, et al.
Published: (2024)
ADMM Algorithms for Residual Network Training: Convergence Analysis and Parallel Implementation
by: Xu, Jintao, et al.
Published: (2023)
by: Xu, Jintao, et al.
Published: (2023)
Convergence of Implicit Gradient Descent for Training Two-Layer Physics-Informed Neural Networks
by: Xu, Xianliang, et al.
Published: (2024)
by: Xu, Xianliang, et al.
Published: (2024)
Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence
by: Liu, Yuxing, et al.
Published: (2025)
by: Liu, Yuxing, et al.
Published: (2025)
Convergence of SGD for Training Neural Networks with Sliced Wasserstein Losses
by: Tanguy, Eloi
Published: (2023)
by: Tanguy, Eloi
Published: (2023)
A Validation Approach to Over-parameterized Matrix and Image Recovery
by: Ding, Lijun, et al.
Published: (2022)
by: Ding, Lijun, et al.
Published: (2022)
Linear Convergence of the Frank-Wolfe Algorithm over Product Polytopes
by: Iommazzo, Gabriele, et al.
Published: (2025)
by: Iommazzo, Gabriele, et al.
Published: (2025)
Linearly Convergent Algorithms for Nonsmooth Problems with Unknown Smooth Pieces
by: Zhang, Zhe, et al.
Published: (2025)
by: Zhang, Zhe, et al.
Published: (2025)
Convergence and Implicit Bias of Gradient Descent on Continual Linear Classification
by: Jung, Hyunji, et al.
Published: (2025)
by: Jung, Hyunji, et al.
Published: (2025)
An Improved Finite-time Analysis of Temporal Difference Learning with Deep Neural Networks
by: Ke, Zhifa, et al.
Published: (2024)
by: Ke, Zhifa, et al.
Published: (2024)
Sobolev Gradient Ascent for Optimal Transport: Barycenter Optimization and Convergence Analysis
by: Kim, Kaheon, et al.
Published: (2025)
by: Kim, Kaheon, et al.
Published: (2025)
Weak Convergence Analysis of Online Neural Actor-Critic Algorithms
by: Lam, Samuel Chun-Hei, et al.
Published: (2024)
by: Lam, Samuel Chun-Hei, et al.
Published: (2024)
Similar Items
-
On the Convergence of Gradient Descent on Learning Transformers with Residual Connections
by: Qin, Zhen, et al.
Published: (2025) -
Learning to Adapt: In-Context Learning Beyond Stationarity
by: Qin, Zhen, et al.
Published: (2026) -
Robust Low-rank Tensor Train Recovery
by: Qin, Zhen, et al.
Published: (2024) -
Computational and Statistical Guarantees for Tensor-on-Tensor Regression with Tensor Train Decomposition
by: Qin, Zhen, et al.
Published: (2024) -
Data Collaboration Analysis with Orthonormal Basis Selection and Alignment
by: Nosaka, Keiyu, et al.
Published: (2024)