The Polar Express: Optimal Matrix Sign Methods and Their Application to the Muon Algorithm
Fuente:
arXiv
Saved in:
| Main Authors: | Amsel, Noah, Persson, David, Musco, Christopher, Gower, Robert M. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Data-induced multiscale losses and efficient multirate gradient descent schemes
by: He, Juncai, et al.
Published: (2024)
by: He, Juncai, et al.
Published: (2024)
When can a neural operator replace a coarse solve? Architectural principles for two-level preconditioning
by: Melchers, Hugo, et al.
Published: (2026)
by: Melchers, Hugo, et al.
Published: (2026)
Application of Sensitivity Analysis Methods for Studying Neural Network Models
by: Miao, Jiaxuan, et al.
Published: (2025)
by: Miao, Jiaxuan, et al.
Published: (2025)
Deep Neural Networks with General Activations: Super-Convergence in Sobolev Norms
by: Yang, Yahong, et al.
Published: (2025)
by: Yang, Yahong, et al.
Published: (2025)
Learning Hamiltonian flows from numerical integrators and examples
by: Fang, Rui, et al.
Published: (2025)
by: Fang, Rui, et al.
Published: (2025)
Alternately-optimized SNN method for acoustic scattering problem in unbounded domain
by: Song, Haoming, et al.
Published: (2025)
by: Song, Haoming, et al.
Published: (2025)
Numerical Solution of Mixed-Dimensional PDEs Using a Neural Preconditioner
by: Dimola, Nunzio, et al.
Published: (2025)
by: Dimola, Nunzio, et al.
Published: (2025)
Challenges in automatic differentiation and numerical integration in physics-informed neural networks modelling
by: Daněk, Josef, et al.
Published: (2024)
by: Daněk, Josef, et al.
Published: (2024)
RANDSMAPs: Random-Feature/multi-Scale Neural Decoders with Mass Preservation
by: Patsatzis, Dimitrios G., et al.
Published: (2026)
by: Patsatzis, Dimitrios G., et al.
Published: (2026)
Sprecher Networks: A Parameter-Efficient Kolmogorov-Arnold Architecture
by: Hägg, Christian, et al.
Published: (2025)
by: Hägg, Christian, et al.
Published: (2025)
Primal-Dual Sample Complexity Bounds for Constrained Markov Decision Processes with Multiple Constraints
by: Buckley, Max, et al.
Published: (2025)
by: Buckley, Max, et al.
Published: (2025)
Quantum Deep Learning Still Needs a Quantum Leap
by: Gundlach, Hans, et al.
Published: (2025)
by: Gundlach, Hans, et al.
Published: (2025)
Optimal Linear Baseline Models for Scientific Machine Learning
by: DeLise, Alexander, et al.
Published: (2025)
by: DeLise, Alexander, et al.
Published: (2025)
On the Convergence Behavior of Preconditioned Gradient Descent Toward the Rich Learning Regime
by: Jiang, Shuai, et al.
Published: (2026)
by: Jiang, Shuai, et al.
Published: (2026)
Sparse Training of Neural Networks based on Multilevel Mirror Descent
by: Lunk, Yannick, et al.
Published: (2026)
by: Lunk, Yannick, et al.
Published: (2026)
NeurOptimisation: The Spiking Way to Evolve
by: Cruz-Duarte, Jorge Mario, et al.
Published: (2025)
by: Cruz-Duarte, Jorge Mario, et al.
Published: (2025)
Polyharmonic Cascade
by: Bakhvalov, Yuriy N.
Published: (2025)
by: Bakhvalov, Yuriy N.
Published: (2025)
On the study of frequency control and spectral bias in Wavelet-Based Kolmogorov Arnold networks: A path to physics-informed KANs
by: Meshir, Juan Daniel, et al.
Published: (2025)
by: Meshir, Juan Daniel, et al.
Published: (2025)
Neural Green's Operators for Parametric Partial Differential Equations
by: Melchers, Hugo, et al.
Published: (2024)
by: Melchers, Hugo, et al.
Published: (2024)
Physics Informed Differentiable Solvers for Learning Parametric Solution Manifolds in Heterogeneous Physical Systems
by: Panahi, Milad, et al.
Published: (2026)
by: Panahi, Milad, et al.
Published: (2026)
Total Generalized Variation regularization closes the gap between neural-eld and classical methods in seismic travel-time tomography
by: Kurosawa, Isao
Published: (2026)
by: Kurosawa, Isao
Published: (2026)
Structured Multidimensional Representation Learning for Large Language Models
by: Ichi, Alaa El, et al.
Published: (2026)
by: Ichi, Alaa El, et al.
Published: (2026)
Deceptron: Learned Local Inverses for Fast and Stable Physics Inversion
by: Kachhadiya, Aaditya L.
Published: (2025)
by: Kachhadiya, Aaditya L.
Published: (2025)
Frequency Bias and OOD Generalization in Neural Operators under a Variable-Coefficient Wave Equation
by: Xie, Runlong, et al.
Published: (2026)
by: Xie, Runlong, et al.
Published: (2026)
Non-convergence to the optimal risk for Adam and stochastic gradient descent optimization in the training of deep neural networks
by: Do, Thang, et al.
Published: (2025)
by: Do, Thang, et al.
Published: (2025)
Learning Nonlinear Finite Element Solution Operators using Multilayer Perceptrons and Energy Minimization
by: Larson, Mats G., et al.
Published: (2024)
by: Larson, Mats G., et al.
Published: (2024)
Kourkoutas-Beta: A Sunspike-Driven Adam Optimizer with Desert Flair
by: Kassinos, Stavros C.
Published: (2025)
by: Kassinos, Stavros C.
Published: (2025)
Adam Improves Muon: Adaptive Moment Estimation with Orthogonalized Momentum
by: Zhang, Minxin, et al.
Published: (2026)
by: Zhang, Minxin, et al.
Published: (2026)
Comparing EPGP Surrogates and Finite Elements Under Degree-of-Freedom Parity
by: Amo, Obed, et al.
Published: (2025)
by: Amo, Obed, et al.
Published: (2025)
A Classical-Quantum Hybrid Architecture for Physics-Informed Neural Networks
by: Lantigua, Said, et al.
Published: (2025)
by: Lantigua, Said, et al.
Published: (2025)
Learning Chaotic Dynamics through Second-Order Geometric Supervision
by: Kang, Shinhoo, et al.
Published: (2026)
by: Kang, Shinhoo, et al.
Published: (2026)
Nearly Optimal Approximation of Matrix Functions by the Lanczos Method
by: Amsel, Noah, et al.
Published: (2023)
by: Amsel, Noah, et al.
Published: (2023)
Local properties of neural networks through the lens of layer-wise Hessians
by: Bolshim, Maxim, et al.
Published: (2025)
by: Bolshim, Maxim, et al.
Published: (2025)
ZetA: A Riemann Zeta-Scaled Extension of Adam for Deep Learning
by: BC, Samiksha
Published: (2025)
by: BC, Samiksha
Published: (2025)
Inter-Layer Hessian Analysis of Neural Networks with DAG Architectures
by: Bolshim, Maxim, et al.
Published: (2026)
by: Bolshim, Maxim, et al.
Published: (2026)
Stochastic Estimation of the Layer-wise Hessian Trace for Monitoring Neural-network Training
by: Bolshim, Maxim, et al.
Published: (2026)
by: Bolshim, Maxim, et al.
Published: (2026)
Diagnosing Failure Modes of Neural Operators Across Diverse PDE Families
by: Shikhman, Lennon
Published: (2026)
by: Shikhman, Lennon
Published: (2026)
A deep solver for backward stochastic Volterra integral equations
by: Andersson, Kristoffer, et al.
Published: (2025)
by: Andersson, Kristoffer, et al.
Published: (2025)
Differentiable Radar Ambiguity Functions: Mathematical Formulation and Computational Implementation
by: Iniesta, Marc Bara
Published: (2025)
by: Iniesta, Marc Bara
Published: (2025)
Understanding the Nature of Generative AI as Threshold Logic in High-Dimensional Space
by: Levin, Ilya
Published: (2026)
by: Levin, Ilya
Published: (2026)
Similar Items
-
Data-induced multiscale losses and efficient multirate gradient descent schemes
by: He, Juncai, et al.
Published: (2024) -
When can a neural operator replace a coarse solve? Architectural principles for two-level preconditioning
by: Melchers, Hugo, et al.
Published: (2026) -
Application of Sensitivity Analysis Methods for Studying Neural Network Models
by: Miao, Jiaxuan, et al.
Published: (2025) -
Deep Neural Networks with General Activations: Super-Convergence in Sobolev Norms
by: Yang, Yahong, et al.
Published: (2025) -
Learning Hamiltonian flows from numerical integrators and examples
by: Fang, Rui, et al.
Published: (2025)