Saved in:
| Main Author: | Thomas, Stephen J. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2605.03109 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cascade Token Selection for Transformer Attention Acceleration
by: Thomas, Stephen J.
Published: (2026)
by: Thomas, Stephen J.
Published: (2026)
SVD-Preconditioned Gradient Descent Method for Solving Nonlinear Least Squares Problems
by: Chang, Zhipeng, et al.
Published: (2026)
by: Chang, Zhipeng, et al.
Published: (2026)
A Neural-Operator Preconditioned Newton Method for Accelerated Nonlinear Solvers
by: Lee, Youngkyu, et al.
Published: (2025)
by: Lee, Youngkyu, et al.
Published: (2025)
Faster Adaptive Optimization via Expected Gradient Outer Product Reparameterization
by: DePavia, Adela, et al.
Published: (2025)
by: DePavia, Adela, et al.
Published: (2025)
A Layer Separation Optimization Framework for Cross-Entropy Training in Deep Learning
by: Liu, Yaru, et al.
Published: (2026)
by: Liu, Yaru, et al.
Published: (2026)
Holonorm
by: Yongueng, Daryl Noupa, et al.
Published: (2025)
by: Yongueng, Daryl Noupa, et al.
Published: (2025)
Objective Value Change and Shape-Based Accelerated Optimization for the Neural Network Approximation
by: Xie, Pengcheng, et al.
Published: (2025)
by: Xie, Pengcheng, et al.
Published: (2025)
Progressive Power Homotopy for Non-convex Optimization
by: Xu, Chen
Published: (2026)
by: Xu, Chen
Published: (2026)
Sample-wise Constrained Learning via a Sequential Penalty Approach with Applications in Image Processing
by: Lanzillotta, Francesca, et al.
Published: (2026)
by: Lanzillotta, Francesca, et al.
Published: (2026)
Global Optimization with A Power-Transformed Objective and Gaussian Smoothing
by: Xu, Chen
Published: (2024)
by: Xu, Chen
Published: (2024)
BROS: Bias-Corrected Randomized Subspaces for Memory-Efficient Single-Loop Bilevel Optimization
by: Zhang, Hengrui, et al.
Published: (2026)
by: Zhang, Hengrui, et al.
Published: (2026)
Effectively Leveraging Momentum Terms in Stochastic Line Search Frameworks for Fast Optimization of Finite-Sum Problems
by: Lapucci, Matteo, et al.
Published: (2024)
by: Lapucci, Matteo, et al.
Published: (2024)
Convergence Conditions for Stochastic Line Search Based Optimization of Over-parametrized Models
by: Lapucci, Matteo, et al.
Published: (2024)
by: Lapucci, Matteo, et al.
Published: (2024)
Power Homotopy for Zeroth-Order Non-Convex Optimizations
by: Xu, Chen
Published: (2025)
by: Xu, Chen
Published: (2025)
A Hybrid Iterative Neural Solver Based on Spectral Analysis for Parametric PDEs
by: Cui, Chen, et al.
Published: (2024)
by: Cui, Chen, et al.
Published: (2024)
Error Bound Analysis for the Regularized Loss of Deep Linear Neural Networks
by: Chen, Po, et al.
Published: (2025)
by: Chen, Po, et al.
Published: (2025)
Refining Graphical Neural Network Predictions Using Flow Matching for Optimal Power Flow with Constraint-Satisfaction Guarantee
by: Khanal, Kshitiz
Published: (2025)
by: Khanal, Kshitiz
Published: (2025)
Inter-Layer Hessian Analysis of Neural Networks with DAG Architectures
by: Bolshim, Maxim, et al.
Published: (2026)
by: Bolshim, Maxim, et al.
Published: (2026)
Stochastic Estimation of the Layer-wise Hessian Trace for Monitoring Neural-network Training
by: Bolshim, Maxim, et al.
Published: (2026)
by: Bolshim, Maxim, et al.
Published: (2026)
Local properties of neural networks through the lens of layer-wise Hessians
by: Bolshim, Maxim, et al.
Published: (2025)
by: Bolshim, Maxim, et al.
Published: (2025)
Towards Solving Polynomial-Objective Integer Programming with Hypergraph Neural Networks
by: Li, Minshuo, et al.
Published: (2026)
by: Li, Minshuo, et al.
Published: (2026)
Two-level overlapping additive Schwarz preconditioner for training scientific machine learning applications
by: Lee, Youngkyu, et al.
Published: (2024)
by: Lee, Youngkyu, et al.
Published: (2024)
Enhancing training of physics-informed neural networks using domain-decomposition based preconditioning strategies
by: Kopaničáková, Alena, et al.
Published: (2023)
by: Kopaničáková, Alena, et al.
Published: (2023)
Hybrid Least Squares/Gradient Descent Methods for DeepONets
by: Choi, Jun, et al.
Published: (2025)
by: Choi, Jun, et al.
Published: (2025)
Generalizing Adam to Manifolds for Efficiently Training Transformers
by: Brantner, Benedikt
Published: (2023)
by: Brantner, Benedikt
Published: (2023)
On the fast convergence of minibatch heavy ball momentum
by: Bollapragada, Raghu, et al.
Published: (2022)
by: Bollapragada, Raghu, et al.
Published: (2022)
Deep Adaptive Dimension Reduction for Bayesian Inference in Inverse Problems
by: Wang, Yueyang, et al.
Published: (2026)
by: Wang, Yueyang, et al.
Published: (2026)
An Augmented Lagrangian Method for Training Recurrent Neural Networks
by: Wang, Yue, et al.
Published: (2024)
by: Wang, Yue, et al.
Published: (2024)
Preconditioned subgradient method for composite optimization: overparameterization and fast convergence
by: Díaz, Mateo, et al.
Published: (2025)
by: Díaz, Mateo, et al.
Published: (2025)
Edge-Wise Graph-Instructed Neural Networks
by: Della Santa, Francesco, et al.
Published: (2024)
by: Della Santa, Francesco, et al.
Published: (2024)
A Paired Autoencoder Framework for Inverse Problems via Bayes Risk Minimization
by: Hart, Emma, et al.
Published: (2025)
by: Hart, Emma, et al.
Published: (2025)
Randomized Matrix Sketching for Neural Network Training and Gradient Monitoring
by: Antil, Harbir, et al.
Published: (2025)
by: Antil, Harbir, et al.
Published: (2025)
Iterative Methods for Full-Scale Gaussian Process Approximations for Large Spatial Data
by: Gyger, Tim, et al.
Published: (2024)
by: Gyger, Tim, et al.
Published: (2024)
Residual Multi-Fidelity Neural Network Computing
by: Davis, Owen, et al.
Published: (2023)
by: Davis, Owen, et al.
Published: (2023)
Beyond Least Squares: Robust Regression Transformer (R2T)
by: Gutierrez, Roman, et al.
Published: (2025)
by: Gutierrez, Roman, et al.
Published: (2025)
Volume-Preserving Transformers for Learning Time Series Data with Structure
by: Brantner, Benedikt, et al.
Published: (2023)
by: Brantner, Benedikt, et al.
Published: (2023)
Neural Preconditioning via Krylov Subspace Geometry
by: Dimola, Nunzio, et al.
Published: (2025)
by: Dimola, Nunzio, et al.
Published: (2025)
CAO: Curvature-Adaptive Optimization via Periodic Low-Rank Hessian Sketching
by: Du, Wenzhang
Published: (2025)
by: Du, Wenzhang
Published: (2025)
Have ASkotch: A Neat Solution for Large-scale Kernel Ridge Regression
by: Rathore, Pratik, et al.
Published: (2024)
by: Rathore, Pratik, et al.
Published: (2024)
Convergence of gradient descent for deep neural networks
by: Chatterjee, Sourav
Published: (2022)
by: Chatterjee, Sourav
Published: (2022)
Similar Items
-
Cascade Token Selection for Transformer Attention Acceleration
by: Thomas, Stephen J.
Published: (2026) -
SVD-Preconditioned Gradient Descent Method for Solving Nonlinear Least Squares Problems
by: Chang, Zhipeng, et al.
Published: (2026) -
A Neural-Operator Preconditioned Newton Method for Accelerated Nonlinear Solvers
by: Lee, Youngkyu, et al.
Published: (2025) -
Faster Adaptive Optimization via Expected Gradient Outer Product Reparameterization
by: DePavia, Adela, et al.
Published: (2025) -
A Layer Separation Optimization Framework for Cross-Entropy Training in Deep Learning
by: Liu, Yaru, et al.
Published: (2026)