Adam Improves Muon: Adaptive Moment Estimation with Orthogonalized Momentum
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Minxin, Liu, Yuxuan, Schaeffer, Hayden |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Kourkoutas-Beta: A Sunspike-Driven Adam Optimizer with Desert Flair
von: Kassinos, Stavros C.
Veröffentlicht: (2025)
von: Kassinos, Stavros C.
Veröffentlicht: (2025)
On the Convergence Behavior of Preconditioned Gradient Descent Toward the Rich Learning Regime
von: Jiang, Shuai, et al.
Veröffentlicht: (2026)
von: Jiang, Shuai, et al.
Veröffentlicht: (2026)
Multi-Hypothesis Prediction for Portfolio Optimization: A Structured Ensemble Learning Approach to Risk Diversification
von: Dominguez, Alejandro Rodriguez, et al.
Veröffentlicht: (2025)
von: Dominguez, Alejandro Rodriguez, et al.
Veröffentlicht: (2025)
ZetA: A Riemann Zeta-Scaled Extension of Adam for Deep Learning
von: BC, Samiksha
Veröffentlicht: (2025)
von: BC, Samiksha
Veröffentlicht: (2025)
Stochastic Estimation of the Layer-wise Hessian Trace for Monitoring Neural-network Training
von: Bolshim, Maxim, et al.
Veröffentlicht: (2026)
von: Bolshim, Maxim, et al.
Veröffentlicht: (2026)
On-Average Stability of Multipass Preconditioned SGD and Effective Dimension
von: Vary, Simon, et al.
Veröffentlicht: (2026)
von: Vary, Simon, et al.
Veröffentlicht: (2026)
FlowAdam: Implicit Regularization via Geometry-Aware Soft Momentum Injection
von: Singh, Devender, et al.
Veröffentlicht: (2026)
von: Singh, Devender, et al.
Veröffentlicht: (2026)
CAO: Curvature-Adaptive Optimization via Periodic Low-Rank Hessian Sketching
von: Du, Wenzhang
Veröffentlicht: (2025)
von: Du, Wenzhang
Veröffentlicht: (2025)
IPAS: An Adaptive Sample Size Method for Weighted Finite Sum Problems with Linear Equality Constraints
von: Krejić, Nataša, et al.
Veröffentlicht: (2025)
von: Krejić, Nataša, et al.
Veröffentlicht: (2025)
Local properties of neural networks through the lens of layer-wise Hessians
von: Bolshim, Maxim, et al.
Veröffentlicht: (2025)
von: Bolshim, Maxim, et al.
Veröffentlicht: (2025)
Inter-Layer Hessian Analysis of Neural Networks with DAG Architectures
von: Bolshim, Maxim, et al.
Veröffentlicht: (2026)
von: Bolshim, Maxim, et al.
Veröffentlicht: (2026)
Fractional-Boundary-Regularized Deep Galerkin Method for Variational Inequalities in Mixed Optimal Stopping and Control
von: Zhao, Yun, et al.
Veröffentlicht: (2025)
von: Zhao, Yun, et al.
Veröffentlicht: (2025)
Machine Collaboration
von: Liu, Qingfeng, et al.
Veröffentlicht: (2021)
von: Liu, Qingfeng, et al.
Veröffentlicht: (2021)
Multiple data-driven missing imputation
von: Kavun, Sergii
Veröffentlicht: (2025)
von: Kavun, Sergii
Veröffentlicht: (2025)
Refining Graphical Neural Network Predictions Using Flow Matching for Optimal Power Flow with Constraint-Satisfaction Guarantee
von: Khanal, Kshitiz
Veröffentlicht: (2025)
von: Khanal, Kshitiz
Veröffentlicht: (2025)
i-DEQ: A stable inertial deep equilibrium model for image restoration
von: Clerc, Antonin, et al.
Veröffentlicht: (2026)
von: Clerc, Antonin, et al.
Veröffentlicht: (2026)
ASMOP: Additional sampling stochastic trust region method for multi-objective problems
von: Jerinkić, Nataša Krklec, et al.
Veröffentlicht: (2025)
von: Jerinkić, Nataša Krklec, et al.
Veröffentlicht: (2025)
Machine Learning Algorithms for Improving Black Box Optimization Solvers
von: Kimiaei, Morteza, et al.
Veröffentlicht: (2025)
von: Kimiaei, Morteza, et al.
Veröffentlicht: (2025)
Autoencoders in Function Space
von: Bunker, Justin, et al.
Veröffentlicht: (2024)
von: Bunker, Justin, et al.
Veröffentlicht: (2024)
SAGA: A Sequence-Adaptive Generative Architecture for Multi-Horizon Probabilistic Forecasting with Adaptive Temporal Conformal Prediction
von: Lundström-Imanov, Gustav Olaf Yunus Laitinen-Fredriksson, et al.
Veröffentlicht: (2026)
von: Lundström-Imanov, Gustav Olaf Yunus Laitinen-Fredriksson, et al.
Veröffentlicht: (2026)
$δ$-STEAL: LLM Stealing Attack with Local Differential Privacy
von: Dang, Kieu, et al.
Veröffentlicht: (2025)
von: Dang, Kieu, et al.
Veröffentlicht: (2025)
Deep Legendre Transform
von: Minabutdinov, Aleksey, et al.
Veröffentlicht: (2025)
von: Minabutdinov, Aleksey, et al.
Veröffentlicht: (2025)
Deceptron: Learned Local Inverses for Fast and Stable Physics Inversion
von: Kachhadiya, Aaditya L.
Veröffentlicht: (2025)
von: Kachhadiya, Aaditya L.
Veröffentlicht: (2025)
G-Sim: Generative Simulations with Large Language Models and Gradient-Free Calibration
von: Holt, Samuel, et al.
Veröffentlicht: (2025)
von: Holt, Samuel, et al.
Veröffentlicht: (2025)
Constant-Target Energy Matching: A Unified Framework for Continuous and Discrete Density Estimation
von: Zeng, Zhijun, et al.
Veröffentlicht: (2026)
von: Zeng, Zhijun, et al.
Veröffentlicht: (2026)
Benchmarking Generative AI Against Bayesian Optimization for Constrained Multi-Objective Inverse Design
von: Awan, Muhammad Bilal, et al.
Veröffentlicht: (2025)
von: Awan, Muhammad Bilal, et al.
Veröffentlicht: (2025)
The Non-Linearity Perturbation Threshold: Width Scaling and Landscape Bifurcations in Deep Learning
von: Alexander, Michael
Veröffentlicht: (2026)
von: Alexander, Michael
Veröffentlicht: (2026)
Data-induced multiscale losses and efficient multirate gradient descent schemes
von: He, Juncai, et al.
Veröffentlicht: (2024)
von: He, Juncai, et al.
Veröffentlicht: (2024)
The Neural Differential Manifold: An Architecture with Explicit Geometric Structure
von: Zhang, Di
Veröffentlicht: (2025)
von: Zhang, Di
Veröffentlicht: (2025)
DISCOVER: A Physics-Informed, GPU-Accelerated Symbolic Regression Framework
von: Gajera, Udaykumar, et al.
Veröffentlicht: (2026)
von: Gajera, Udaykumar, et al.
Veröffentlicht: (2026)
Explainable Attention-Based LSTM Framework for Early Detection of AI-Assisted Ransomware via File System Behavioral Analysis
von: Nayak, Prabhudarshi, et al.
Veröffentlicht: (2026)
von: Nayak, Prabhudarshi, et al.
Veröffentlicht: (2026)
On the study of frequency control and spectral bias in Wavelet-Based Kolmogorov Arnold networks: A path to physics-informed KANs
von: Meshir, Juan Daniel, et al.
Veröffentlicht: (2025)
von: Meshir, Juan Daniel, et al.
Veröffentlicht: (2025)
NeurOptimisation: The Spiking Way to Evolve
von: Cruz-Duarte, Jorge Mario, et al.
Veröffentlicht: (2025)
von: Cruz-Duarte, Jorge Mario, et al.
Veröffentlicht: (2025)
The Polar Express: Optimal Matrix Sign Methods and Their Application to the Muon Algorithm
von: Amsel, Noah, et al.
Veröffentlicht: (2025)
von: Amsel, Noah, et al.
Veröffentlicht: (2025)
"Abuse Risks are Often Inherent to Product Features": Exploring AI Vendors' Bug Bounty and Responsible Disclosure Policies
von: Piao, Yangheran, et al.
Veröffentlicht: (2025)
von: Piao, Yangheran, et al.
Veröffentlicht: (2025)
TED++: Submanifold-Aware Backdoor Detection via Layerwise Tubular-Neighbourhood Screening
von: Le, Nam, et al.
Veröffentlicht: (2025)
von: Le, Nam, et al.
Veröffentlicht: (2025)
Differentiable Optimization Layers for Guaranteed Fairness in Deep Learning
von: Troxell, David, et al.
Veröffentlicht: (2026)
von: Troxell, David, et al.
Veröffentlicht: (2026)
EB-gMCR: Energy-Based Generative Modeling for Signal Unmixing and Multivariate Curve Resolution
von: Chang, Yu-Tang, et al.
Veröffentlicht: (2025)
von: Chang, Yu-Tang, et al.
Veröffentlicht: (2025)
Sparse Training of Neural Networks based on Multilevel Mirror Descent
von: Lunk, Yannick, et al.
Veröffentlicht: (2026)
von: Lunk, Yannick, et al.
Veröffentlicht: (2026)
Gradient descent provably escapes saddle points in the training of shallow ReLU networks
von: Cheridito, Patrick, et al.
Veröffentlicht: (2022)
von: Cheridito, Patrick, et al.
Veröffentlicht: (2022)
Ähnliche Einträge
-
Kourkoutas-Beta: A Sunspike-Driven Adam Optimizer with Desert Flair
von: Kassinos, Stavros C.
Veröffentlicht: (2025) -
On the Convergence Behavior of Preconditioned Gradient Descent Toward the Rich Learning Regime
von: Jiang, Shuai, et al.
Veröffentlicht: (2026) -
Multi-Hypothesis Prediction for Portfolio Optimization: A Structured Ensemble Learning Approach to Risk Diversification
von: Dominguez, Alejandro Rodriguez, et al.
Veröffentlicht: (2025) -
ZetA: A Riemann Zeta-Scaled Extension of Adam for Deep Learning
von: BC, Samiksha
Veröffentlicht: (2025) -
Stochastic Estimation of the Layer-wise Hessian Trace for Monitoring Neural-network Training
von: Bolshim, Maxim, et al.
Veröffentlicht: (2026)