A Layer Separation Optimization Framework for Cross-Entropy Training in Deep Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Yaru, Ng, Michael K., Gu, Yiqi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Faster Adaptive Optimization via Expected Gradient Outer Product Reparameterization
by: DePavia, Adela, et al.
Published: (2025)
by: DePavia, Adela, et al.
Published: (2025)
An Augmented Lagrangian Method for Training Recurrent Neural Networks
by: Wang, Yue, et al.
Published: (2024)
by: Wang, Yue, et al.
Published: (2024)
Error Bound Analysis for the Regularized Loss of Deep Linear Neural Networks
by: Chen, Po, et al.
Published: (2025)
by: Chen, Po, et al.
Published: (2025)
Effectively Leveraging Momentum Terms in Stochastic Line Search Frameworks for Fast Optimization of Finite-Sum Problems
by: Lapucci, Matteo, et al.
Published: (2024)
by: Lapucci, Matteo, et al.
Published: (2024)
Power Homotopy for Zeroth-Order Non-Convex Optimizations
by: Xu, Chen
Published: (2025)
by: Xu, Chen
Published: (2025)
SVD-Preconditioned Gradient Descent Method for Solving Nonlinear Least Squares Problems
by: Chang, Zhipeng, et al.
Published: (2026)
by: Chang, Zhipeng, et al.
Published: (2026)
Global Optimization with A Power-Transformed Objective and Gaussian Smoothing
by: Xu, Chen
Published: (2024)
by: Xu, Chen
Published: (2024)
Convergence Conditions for Stochastic Line Search Based Optimization of Over-parametrized Models
by: Lapucci, Matteo, et al.
Published: (2024)
by: Lapucci, Matteo, et al.
Published: (2024)
Sample-wise Constrained Learning via a Sequential Penalty Approach with Applications in Image Processing
by: Lanzillotta, Francesca, et al.
Published: (2026)
by: Lanzillotta, Francesca, et al.
Published: (2026)
Progressive Power Homotopy for Non-convex Optimization
by: Xu, Chen
Published: (2026)
by: Xu, Chen
Published: (2026)
Preconditioned subgradient method for composite optimization: overparameterization and fast convergence
by: Díaz, Mateo, et al.
Published: (2025)
by: Díaz, Mateo, et al.
Published: (2025)
Binno: A 1st-order method for Bi-level Nonconvex Nonsmooth Optimization for Matrix Factorizations
by: Selicato, Laura, et al.
Published: (2025)
by: Selicato, Laura, et al.
Published: (2025)
Inexact Riemannian Gradient Descent Method for Nonconvex Optimization
by: Zhou, Juan, et al.
Published: (2024)
by: Zhou, Juan, et al.
Published: (2024)
Parameter-Free Accelerated Quasi-Newton Method for Nonconvex Optimization
by: Marumo, Naoki
Published: (2025)
by: Marumo, Naoki
Published: (2025)
A Unified Zeroth-Order Proximal Newton-Type Framework for Composite Optimization
by: Liu, Zekun, et al.
Published: (2026)
by: Liu, Zekun, et al.
Published: (2026)
Stochastic optimization over proximally smooth sets
by: Davis, Damek, et al.
Published: (2020)
by: Davis, Damek, et al.
Published: (2020)
Technical results on the convergence of quasi-Newton methods for nonsmooth optimization
by: Gebken, Bennet
Published: (2025)
by: Gebken, Bennet
Published: (2025)
A Proximal-Gradient Method for Constrained Optimization
by: Dai, Yutong, et al.
Published: (2024)
by: Dai, Yutong, et al.
Published: (2024)
The Challenges of Optimization For Data Science
by: Varner, Christian, et al.
Published: (2024)
by: Varner, Christian, et al.
Published: (2024)
A Proximal-Gradient Method for Solving Regularized Optimization Problems with General Constraints
by: Curtis, Frank E., et al.
Published: (2025)
by: Curtis, Frank E., et al.
Published: (2025)
Solving nonconvex optimization problems via a second order dynamical system with unbounded damping
by: László, Szilárd Csaba
Published: (2025)
by: László, Szilárd Csaba
Published: (2025)
A Bregman ADMM for Bethe variational problem
by: Khoo, Yuehaw, et al.
Published: (2025)
by: Khoo, Yuehaw, et al.
Published: (2025)
Optimization in Theory and Practice
by: Wright, Stephen J.
Published: (2025)
by: Wright, Stephen J.
Published: (2025)
Nonsmooth Projection-Free Optimization with Functional Constraints
by: Asgari, Kamiar, et al.
Published: (2023)
by: Asgari, Kamiar, et al.
Published: (2023)
ASPEN: An Additional Sampling Penalty Method for Finite-Sum Optimization Problems with Nonlinear Equality Constraints
by: Krejić, Nataša, et al.
Published: (2025)
by: Krejić, Nataša, et al.
Published: (2025)
Heavy Ball Momentum for Non-Strongly Convex Optimization
by: Aujol, Jean-François, et al.
Published: (2024)
by: Aujol, Jean-François, et al.
Published: (2024)
Refining Graphical Neural Network Predictions Using Flow Matching for Optimal Power Flow with Constraint-Satisfaction Guarantee
by: Khanal, Kshitiz
Published: (2025)
by: Khanal, Kshitiz
Published: (2025)
Objective Value Change and Shape-Based Accelerated Optimization for the Neural Network Approximation
by: Xie, Pengcheng, et al.
Published: (2025)
by: Xie, Pengcheng, et al.
Published: (2025)
A Two Stepsize SQP Method for Nonlinear Equality Constrained Stochastic Optimization
by: O'Neill, Michael J.
Published: (2024)
by: O'Neill, Michael J.
Published: (2024)
Spectrally Constrained Optimization
by: Garner, Casey, et al.
Published: (2023)
by: Garner, Casey, et al.
Published: (2023)
General Constrained Matrix Optimization
by: Garner, Casey, et al.
Published: (2024)
by: Garner, Casey, et al.
Published: (2024)
Holonorm
by: Yongueng, Daryl Noupa, et al.
Published: (2025)
by: Yongueng, Daryl Noupa, et al.
Published: (2025)
Inexact FPPA for the $\ell_0$ Sparse Regularization Problem
by: Fang, Ronglong, et al.
Published: (2024)
by: Fang, Ronglong, et al.
Published: (2024)
A Novel Gradient Methodology with Economical Objective Function Evaluations for Data Science Applications
by: Varner, Christian, et al.
Published: (2023)
by: Varner, Christian, et al.
Published: (2023)
A Novel First-order Method with Event-driven Objective Evaluations
by: Varner, Christian, et al.
Published: (2025)
by: Varner, Christian, et al.
Published: (2025)
Gradient descent with adaptive stepsize converges (nearly) linearly under fourth-order growth
by: Davis, Damek, et al.
Published: (2024)
by: Davis, Damek, et al.
Published: (2024)
Delayed Feedback in Online Non-Convex Optimization: A Non-Stationary Approach with Applications
by: Lara, Felipe, et al.
Published: (2024)
by: Lara, Felipe, et al.
Published: (2024)
Distributed Gradient-Regularized Newton Method: Scheduled Consensus and O(epsilon^{-1}) Global Iteration Complexity
by: Hu, Wei, et al.
Published: (2026)
by: Hu, Wei, et al.
Published: (2026)
Strong Global Convergence of the Consensus-Based Optimization Algorithm
by: Bonandin, Sabrina, et al.
Published: (2025)
by: Bonandin, Sabrina, et al.
Published: (2025)
Adaptive Inertial Method
by: Long, Han, et al.
Published: (2025)
by: Long, Han, et al.
Published: (2025)
Similar Items
-
Faster Adaptive Optimization via Expected Gradient Outer Product Reparameterization
by: DePavia, Adela, et al.
Published: (2025) -
An Augmented Lagrangian Method for Training Recurrent Neural Networks
by: Wang, Yue, et al.
Published: (2024) -
Error Bound Analysis for the Regularized Loss of Deep Linear Neural Networks
by: Chen, Po, et al.
Published: (2025) -
Effectively Leveraging Momentum Terms in Stochastic Line Search Frameworks for Fast Optimization of Finite-Sum Problems
by: Lapucci, Matteo, et al.
Published: (2024) -
Power Homotopy for Zeroth-Order Non-Convex Optimizations
by: Xu, Chen
Published: (2025)