Efficient and provably convergent end-to-end training of deep neural networks with linear constraints
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Zonglin, Gu, Zhexuan, Yuan, Yancheng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Accelerating RLHF Training with Reward Variance Increase
by: Yang, Zonglin, et al.
Published: (2025)
by: Yang, Zonglin, et al.
Published: (2025)
Recurrent neural networks: vanishing and exploding gradients are not the end of the story
by: Zucchet, Nicolas, et al.
Published: (2024)
by: Zucchet, Nicolas, et al.
Published: (2024)
A multiobjective continuation method to compute the regularization path of deep neural networks
by: Amakor, Augustina C., et al.
Published: (2023)
by: Amakor, Augustina C., et al.
Published: (2023)
A space-decoupling framework for optimization on bounded-rank matrices with orthogonally invariant constraints
by: Yang, Yan, et al.
Published: (2025)
by: Yang, Yan, et al.
Published: (2025)
Precise gradient descent training dynamics for finite-width multi-layer neural networks
by: Han, Qiyang, et al.
Published: (2025)
by: Han, Qiyang, et al.
Published: (2025)
Pinet: Optimizing hard-constrained neural networks with orthogonal projection layers
by: Grontas, Panagiotis D., et al.
Published: (2025)
by: Grontas, Panagiotis D., et al.
Published: (2025)
Explicit neural network classifiers for non-separable data
by: Ewald, Patrícia Muñoz
Published: (2025)
by: Ewald, Patrícia Muñoz
Published: (2025)
Demystifying Manifold Constraints in LLM Pre-training
by: An, Kang, et al.
Published: (2026)
by: An, Kang, et al.
Published: (2026)
A multilevel approach to accelerate the training of Transformers
by: Lauga, Guillaume, et al.
Published: (2025)
by: Lauga, Guillaume, et al.
Published: (2025)
AdLoCo: adaptive batching significantly improves communications efficiency and convergence for Large Language Models
by: Kutuzov, Nikolay, et al.
Published: (2025)
by: Kutuzov, Nikolay, et al.
Published: (2025)
How Does Critical Batch Size Scale in Pre-training?
by: Zhang, Hanlin, et al.
Published: (2024)
by: Zhang, Hanlin, et al.
Published: (2024)
Global Convergence and Rich Feature Learning in $L$-Layer Infinite-Width Neural Networks under $μ$P Parametrization
by: Chen, Zixiang, et al.
Published: (2025)
by: Chen, Zixiang, et al.
Published: (2025)
A second-order-like optimizer with adaptive gradient scaling for deep learning
by: Bolte, Jérôme, et al.
Published: (2024)
by: Bolte, Jérôme, et al.
Published: (2024)
HOT: An Efficient Halpern Accelerating Algorithm for Optimal Transport Problems
by: Zhang, Guojun, et al.
Published: (2024)
by: Zhang, Guojun, et al.
Published: (2024)
Efficient model predictive control for nonlinear systems modelled by deep neural networks
by: Lan, Jianglin
Published: (2024)
by: Lan, Jianglin
Published: (2024)
Data Uniformity Improves Training Efficiency and More, with a Convergence Framework Beyond the NTK Regime
by: Wang, Yuqing, et al.
Published: (2025)
by: Wang, Yuqing, et al.
Published: (2025)
Bilevel reinforcement learning via the development of hyper-gradient without lower-level convexity
by: Yang, Yan, et al.
Published: (2024)
by: Yang, Yan, et al.
Published: (2024)
CLCR: Contrastive Learning-based Constraint Reordering for Efficient MILP Solving
by: Zeng, Shuli, et al.
Published: (2025)
by: Zeng, Shuli, et al.
Published: (2025)
Non-convergence to global minimizers for Adam and stochastic gradient descent optimization and constructions of local minimizers in the training of artificial neural networks
by: Jentzen, Arnulf, et al.
Published: (2024)
by: Jentzen, Arnulf, et al.
Published: (2024)
Double Momentum Method for Lower-Level Constrained Bilevel Optimization
by: Shi, Wanli, et al.
Published: (2024)
by: Shi, Wanli, et al.
Published: (2024)
Learning a local trading strategy: deep reinforcement learning for grid-scale renewable energy integration
by: Ju, Caleb, et al.
Published: (2024)
by: Ju, Caleb, et al.
Published: (2024)
Approximation and interpolation of deep neural networks
by: Constantinescu, Vlad-Raul, et al.
Published: (2023)
by: Constantinescu, Vlad-Raul, et al.
Published: (2023)
Wasserstein distributional adversarial training for deep neural networks
by: Bai, Xingjian, et al.
Published: (2025)
by: Bai, Xingjian, et al.
Published: (2025)
Centrality-Based Pruning for Efficient Echo State Networks
by: Laudari, Sudip
Published: (2026)
by: Laudari, Sudip
Published: (2026)
GANQ: GPU-Adaptive Non-Uniform Quantization for Large Language Models
by: Zhao, Pengxiang, et al.
Published: (2025)
by: Zhao, Pengxiang, et al.
Published: (2025)
SPAP: Structured Pruning via Alternating Optimization and Penalty Methods
by: Hu, Hanyu, et al.
Published: (2025)
by: Hu, Hanyu, et al.
Published: (2025)
Towards Efficient Constraint Handling in Neural Solvers for Routing Problems
by: Bi, Jieyi, et al.
Published: (2026)
by: Bi, Jieyi, et al.
Published: (2026)
Memory-Efficient LLM Pretraining via Minimalist Optimizer Design
by: Glentis, Athanasios, et al.
Published: (2025)
by: Glentis, Athanasios, et al.
Published: (2025)
A Theoretical Framework for Auxiliary-Loss-Free Load Balancing of Sparse Mixture-of-Experts in Large-Scale AI Models
by: Han, X. Y., et al.
Published: (2025)
by: Han, X. Y., et al.
Published: (2025)
Scaling physics-informed hard constraints with mixture-of-experts
by: Chalapathi, Nithin, et al.
Published: (2024)
by: Chalapathi, Nithin, et al.
Published: (2024)
ECPv2: Fast, Efficient, and Scalable Global Optimization of Lipschitz Functions
by: Fourati, Fares, et al.
Published: (2025)
by: Fourati, Fares, et al.
Published: (2025)
Ginger: An Efficient Curvature Approximation with Linear Complexity for General Neural Networks
by: Hao, Yongchang, et al.
Published: (2024)
by: Hao, Yongchang, et al.
Published: (2024)
Towards Efficient Risk-Sensitive Policy Gradient: An Iteration Complexity Analysis
by: Liu, Rui, et al.
Published: (2024)
by: Liu, Rui, et al.
Published: (2024)
Deep Reinforcement Learning for Traveling Purchaser Problems
by: Yuan, Haofeng, et al.
Published: (2024)
by: Yuan, Haofeng, et al.
Published: (2024)
An Efficient Learning-based Solver Comparable to Metaheuristics for the Capacitated Arc Routing Problem
by: Guo, Runze, et al.
Published: (2024)
by: Guo, Runze, et al.
Published: (2024)
Data-driven Projection Generation for Efficiently Solving Heterogeneous Quadratic Programming Problems
by: Iwata, Tomoharu, et al.
Published: (2025)
by: Iwata, Tomoharu, et al.
Published: (2025)
Remove that Square Root: A New Efficient Scale-Invariant Version of AdaGrad
by: Choudhury, Sayantan, et al.
Published: (2024)
by: Choudhury, Sayantan, et al.
Published: (2024)
Towards Faster Decentralized Stochastic Optimization with Communication Compression
by: Islamov, Rustem, et al.
Published: (2024)
by: Islamov, Rustem, et al.
Published: (2024)
Parameter-Efficient Distributional RL via Normalizing Flows and a Geometry-Aware Cramér Surrogate
by: C., Simo Alami, et al.
Published: (2025)
by: C., Simo Alami, et al.
Published: (2025)
Federated Distributionally Robust Optimization with Non-Convex Objectives: Algorithm and Analysis
by: Jiao, Yang, et al.
Published: (2023)
by: Jiao, Yang, et al.
Published: (2023)
Similar Items
-
Accelerating RLHF Training with Reward Variance Increase
by: Yang, Zonglin, et al.
Published: (2025) -
Recurrent neural networks: vanishing and exploding gradients are not the end of the story
by: Zucchet, Nicolas, et al.
Published: (2024) -
A multiobjective continuation method to compute the regularization path of deep neural networks
by: Amakor, Augustina C., et al.
Published: (2023) -
A space-decoupling framework for optimization on bounded-rank matrices with orthogonally invariant constraints
by: Yang, Yan, et al.
Published: (2025) -
Precise gradient descent training dynamics for finite-width multi-layer neural networks
by: Han, Qiyang, et al.
Published: (2025)