Muon Dynamics as a Spectral Wasserstein Flow
Fuente:
arXiv
Enregistré dans:
| Auteur principal: | Peyré, Gabriel |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Optimal and Diffusion Transports in Machine Learning
par: Peyré, Gabriel
Publié: (2025)
par: Peyré, Gabriel
Publié: (2025)
Optimal Transport for Machine Learners
par: Peyré, Gabriel
Publié: (2025)
par: Peyré, Gabriel
Publié: (2025)
Preconditioning Benefits of Spectral Orthogonalization in Muon
par: Ma, Jianhao, et autres
Publié: (2026)
par: Ma, Jianhao, et autres
Publié: (2026)
Training Infinitely Deep and Wide Transformers
par: Barboni, Raphaël, et autres
Publié: (2026)
par: Barboni, Raphaël, et autres
Publié: (2026)
MiMuon: Mixed Muon Optimizer with Improved Generalization for Large Models
par: Huang, Feihu, et autres
Publié: (2026)
par: Huang, Feihu, et autres
Publié: (2026)
The Newton-Muon Optimizer
par: Du, Zhehang, et autres
Publié: (2026)
par: Du, Zhehang, et autres
Publié: (2026)
The Mathematics of Artificial Intelligence
par: Peyré, Gabriel
Publié: (2025)
par: Peyré, Gabriel
Publié: (2025)
Abide by the Law and Follow the Flow: Conservation Laws for Gradient Flows
par: Marcotte, Sibylle, et autres
Publié: (2023)
par: Marcotte, Sibylle, et autres
Publié: (2023)
Robust Sublinear Convergence Rates for Iterative Bregman Projections
par: Peyré, Gabriel
Publié: (2026)
par: Peyré, Gabriel
Publié: (2026)
Optimizer-Induced Mode Connectivity: From AdamW to Muon
par: Zhang, Fangzhao, et autres
Publié: (2026)
par: Zhang, Fangzhao, et autres
Publié: (2026)
Muon Outperforms Adam in Tail-End Associative Memory Learning
par: Wang, Shuche, et autres
Publié: (2025)
par: Wang, Shuche, et autres
Publié: (2025)
Keep the Momentum: Conservation Laws beyond Euclidean Gradient Flows
par: Marcotte, Sibylle, et autres
Publié: (2024)
par: Marcotte, Sibylle, et autres
Publié: (2024)
DeMuon: A Decentralized Muon for Matrix Optimization over Graphs
par: He, Chuan, et autres
Publié: (2025)
par: He, Chuan, et autres
Publié: (2025)
Constrained Sliced Wasserstein Embedding
par: NaderiAlizadeh, Navid, et autres
Publié: (2025)
par: NaderiAlizadeh, Navid, et autres
Publié: (2025)
Muon is Not That Special: Random or Inverted Spectra Work Just as Well
par: Shumaylov, Zakhar, et autres
Publié: (2026)
par: Shumaylov, Zakhar, et autres
Publié: (2026)
Stability of Primal-Dual Gradient Flow Dynamics for Multi-Block Convex Optimization Problems
par: Ozaslan, Ibrahim K., et autres
Publié: (2024)
par: Ozaslan, Ibrahim K., et autres
Publié: (2024)
Anytime Training with Schedule-Free Spectral Optimization
par: Apte, Anuj, et autres
Publié: (2026)
par: Apte, Anuj, et autres
Publié: (2026)
Robust $Q$-learning Algorithm for Markov Decision Processes under Wasserstein Uncertainty
par: Neufeld, Ariel, et autres
Publié: (2022)
par: Neufeld, Ariel, et autres
Publié: (2022)
Primal-Dual Spectral Representation for Off-policy Evaluation
par: Hu, Yang, et autres
Publié: (2024)
par: Hu, Yang, et autres
Publié: (2024)
Optimal Control Operator Perspective and a Neural Adaptive Spectral Method
par: Feng, Mingquan, et autres
Publié: (2024)
par: Feng, Mingquan, et autres
Publié: (2024)
On the global convergence of gradient descent for wide shallow models with bounded nonlinearities
par: Petit, Romain, et autres
Publié: (2026)
par: Petit, Romain, et autres
Publié: (2026)
Understanding Optimization in Deep Learning with Central Flows
par: Cohen, Jeremy M., et autres
Publié: (2024)
par: Cohen, Jeremy M., et autres
Publié: (2024)
Flowing Datasets with Wasserstein over Wasserstein Gradient Flows
par: Bonet, Clément, et autres
Publié: (2025)
par: Bonet, Clément, et autres
Publié: (2025)
A Rod Flow Model for Adam at the Edge of Stability
par: Regis, Eric, et autres
Publié: (2026)
par: Regis, Eric, et autres
Publié: (2026)
Parameter-Efficient Distributional RL via Normalizing Flows and a Geometry-Aware Cramér Surrogate
par: C., Simo Alami, et autres
Publié: (2025)
par: C., Simo Alami, et autres
Publié: (2025)
Implicit Regularization of Gradient Flow on One-Layer Softmax Attention
par: Sheen, Heejune, et autres
Publié: (2024)
par: Sheen, Heejune, et autres
Publié: (2024)
gridfm-datakit-v1: A Python Library for Scalable and Realistic Power Flow and Optimal Power Flow Data Generation
par: Puech, Alban, et autres
Publié: (2025)
par: Puech, Alban, et autres
Publié: (2025)
Rod Flow: A Continuous-Time Model for Gradient Descent at the Edge of Stability
par: Regis, Eric, et autres
Publié: (2026)
par: Regis, Eric, et autres
Publié: (2026)
Muon Optimizes Under Spectral Norm Constraints
par: Chen, Lizhang, et autres
Publié: (2025)
par: Chen, Lizhang, et autres
Publié: (2025)
Understanding the training of infinitely deep and wide ResNets with Conditional Optimal Transport
par: Barboni, Raphaël, et autres
Publié: (2024)
par: Barboni, Raphaël, et autres
Publié: (2024)
Ultra-fast feature learning for the training of two-layer neural networks in the two-timescale regime
par: Barboni, Raphaël, et autres
Publié: (2025)
par: Barboni, Raphaël, et autres
Publié: (2025)
From Optimization to Prediction: Transformer-Based Path-Flow Estimation to the Traffic Assignment Problem
par: Ameli, Mostafa, et autres
Publié: (2025)
par: Ameli, Mostafa, et autres
Publié: (2025)
One Rank at a Time: Cascading Error Dynamics in Sequential Learning
par: Vandchali, Mahtab Alizadeh, et autres
Publié: (2025)
par: Vandchali, Mahtab Alizadeh, et autres
Publié: (2025)
PGLearn -- An Open-Source Learning Toolkit for Optimal Power Flow
par: Klamkin, Michael, et autres
Publié: (2025)
par: Klamkin, Michael, et autres
Publié: (2025)
Dynamic Memory Based Adaptive Optimization
par: Szegedy, Balázs, et autres
Publié: (2024)
par: Szegedy, Balázs, et autres
Publié: (2024)
Joint Problems in Learning Multiple Dynamical Systems
par: Niu, Mengjia, et autres
Publié: (2023)
par: Niu, Mengjia, et autres
Publié: (2023)
The Implicit Curriculum: Learning Dynamics in RL with Verifiable Rewards
par: Huang, Yu, et autres
Publié: (2026)
par: Huang, Yu, et autres
Publié: (2026)
Power Constrained Nonstationary Bandits with Habituation and Recovery Dynamics
par: Li, Fengxu, et autres
Publié: (2025)
par: Li, Fengxu, et autres
Publié: (2025)
Almost Bayesian: The Fractal Dynamics of Stochastic Gradient Descent
par: Hennick, Max, et autres
Publié: (2025)
par: Hennick, Max, et autres
Publié: (2025)
Differentiable Optimization for Deep Learning-Enhanced DC Approximation of AC Optimal Power Flow
par: Rosemberg, Andrew, et autres
Publié: (2025)
par: Rosemberg, Andrew, et autres
Publié: (2025)
Documents similaires
-
Optimal and Diffusion Transports in Machine Learning
par: Peyré, Gabriel
Publié: (2025) -
Optimal Transport for Machine Learners
par: Peyré, Gabriel
Publié: (2025) -
Preconditioning Benefits of Spectral Orthogonalization in Muon
par: Ma, Jianhao, et autres
Publié: (2026) -
Training Infinitely Deep and Wide Transformers
par: Barboni, Raphaël, et autres
Publié: (2026) -
MiMuon: Mixed Muon Optimizer with Improved Generalization for Large Models
par: Huang, Feihu, et autres
Publié: (2026)