Muon is Not That Special: Random or Inverted Spectra Work Just as Well
Fuente:
arXiv
Saved in:
| Main Authors: | Shumaylov, Zakhar, Da Costa, Nathaël, Zaika, Peter, Mucsányi, Bálint, Massucco, Alex, Gelberg, Yoav, Schönlieb, Carola-Bibiane, Gal, Yarin, Hennig, Philipp |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Data-driven approaches to inverse problems
by: Schönlieb, Carola-Bibiane, et al.
Published: (2025)
by: Schönlieb, Carola-Bibiane, et al.
Published: (2025)
Rethinking Approximate Gaussian Inference in Classification
by: Mucsányi, Bálint, et al.
Published: (2025)
by: Mucsányi, Bálint, et al.
Published: (2025)
Geometric Gaussian Approximations of Probability Distributions
by: Da Costa, Nathaël, et al.
Published: (2025)
by: Da Costa, Nathaël, et al.
Published: (2025)
When is a System Discoverable from Data? Discovery Requires Chaos
by: Shumaylov, Zakhar, et al.
Published: (2025)
by: Shumaylov, Zakhar, et al.
Published: (2025)
Generalized Lie Symmetries in Physics-Informed Neural Operators
by: Wang, Amy Xiang, et al.
Published: (2025)
by: Wang, Amy Xiang, et al.
Published: (2025)
Lie Algebra Canonicalization: Equivariant Neural Operators under arbitrary Lie Groups
by: Shumaylov, Zakhar, et al.
Published: (2024)
by: Shumaylov, Zakhar, et al.
Published: (2024)
Weakly Convex Regularisers for Inverse Problems: Convergence of Critical Points and Primal-Dual Optimisation
by: Shumaylov, Zakhar, et al.
Published: (2024)
by: Shumaylov, Zakhar, et al.
Published: (2024)
Score-based Pullback Riemannian Geometry: Extracting the Data Manifold Geometry using Anisotropic Flows
by: Diepeveen, Willem, et al.
Published: (2024)
by: Diepeveen, Willem, et al.
Published: (2024)
Hamiltonian Matching for Symplectic Neural Integrators
by: Canizares, Priscilla, et al.
Published: (2024)
by: Canizares, Priscilla, et al.
Published: (2024)
Symplectic Neural Flows for Modeling and Discovery
by: Canizares, Priscilla, et al.
Published: (2024)
by: Canizares, Priscilla, et al.
Published: (2024)
Deep Network Trainability via Persistent Subspace Orthogonality
by: Massucco, Alex, et al.
Published: (2025)
by: Massucco, Alex, et al.
Published: (2025)
Diffeomorphism-Equivariant Neural Networks
by: Oettinger, Josephine Elisabeth, et al.
Published: (2026)
by: Oettinger, Josephine Elisabeth, et al.
Published: (2026)
Training Transformers for KV Cache Compressibility
by: Gelberg, Yoav, et al.
Published: (2026)
by: Gelberg, Yoav, et al.
Published: (2026)
Debiasing Mini-Batch Quadratics for Applications in Deep Learning
by: Tatzel, Lukas, et al.
Published: (2024)
by: Tatzel, Lukas, et al.
Published: (2024)
The Curse of Recursion: Training on Generated Data Makes Models Forget
by: Shumailov, Ilia, et al.
Published: (2023)
by: Shumailov, Ilia, et al.
Published: (2023)
Adaptive Coordinate Transforms for Neural Operators
by: Liu, Chaoyu, et al.
Published: (2026)
by: Liu, Chaoyu, et al.
Published: (2026)
Benchmarking learned algorithms for computed tomography image reconstruction tasks
by: Kiss, Maximilian B., et al.
Published: (2024)
by: Kiss, Maximilian B., et al.
Published: (2024)
Multi-Headed Transformer Architectures as Time-dependent Wasserstein Gradient Flows
by: Massucco, Alex, et al.
Published: (2026)
by: Massucco, Alex, et al.
Published: (2026)
Benchmarking Uncertainty Disentanglement: Specialized Uncertainties for Specialized Tasks
by: Mucsányi, Bálint, et al.
Published: (2024)
by: Mucsányi, Bálint, et al.
Published: (2024)
SpectraKAN: Conditioning Spectral Operators
by: Cheng, Chun-Wun, et al.
Published: (2026)
by: Cheng, Chun-Wun, et al.
Published: (2026)
Expressivity of Bi-Lipschitz Normalizing Flows: A Score-Based Diffusion Perspective
by: Iske, Meira, et al.
Published: (2026)
by: Iske, Meira, et al.
Published: (2026)
Variational Inference Failures Under Model Symmetries: Permutation Invariant Posteriors for Bayesian Neural Networks
by: Gelberg, Yoav, et al.
Published: (2024)
by: Gelberg, Yoav, et al.
Published: (2024)
Sample Path Regularity of Gaussian Processes from the Covariance Kernel
by: Da Costa, Nathaël, et al.
Published: (2023)
by: Da Costa, Nathaël, et al.
Published: (2023)
Iterative Deployment Improves Planning Skills in LLMs
by: Corrêa, Augusto B., et al.
Published: (2025)
by: Corrêa, Augusto B., et al.
Published: (2025)
Boosting Data-Driven Mirror Descent with Randomization, Equivariance, and Acceleration
by: Tan, Hong Ye, et al.
Published: (2023)
by: Tan, Hong Ye, et al.
Published: (2023)
On Information Geometry and Iterative Optimization in Model Compression: Operator Factorization
by: Shumaylov, Zakhar, et al.
Published: (2025)
by: Shumaylov, Zakhar, et al.
Published: (2025)
Spatiotemporal Graph Neural Network Modelling Perfusion MRI
by: Yan, Ruodan, et al.
Published: (2024)
by: Yan, Ruodan, et al.
Published: (2024)
Inverse Evolution Data Augmentation for Neural PDE Solvers
by: Liu, Chaoyu, et al.
Published: (2025)
by: Liu, Chaoyu, et al.
Published: (2025)
Potential Contrast: Properties, Equivalences, and Generalization to Multiple Classes
by: Peaslee, Wallace, et al.
Published: (2025)
by: Peaslee, Wallace, et al.
Published: (2025)
Generative Unordered Flow for Set-Structured Data Generation
by: Li, Yangming, et al.
Published: (2025)
by: Li, Yangming, et al.
Published: (2025)
Enhanced Denoising and Convergent Regularisation Using Tweedie Scaling
by: Khelifa, Naïl, et al.
Published: (2025)
by: Khelifa, Naïl, et al.
Published: (2025)
Stochastic Primal-Dual Three Operator Splitting Algorithm with Extension to Equivariant Regularization-by-Denoising
by: Tang, Junqi, et al.
Published: (2022)
by: Tang, Junqi, et al.
Published: (2022)
Continuous Learned Primal Dual
by: Runkel, Christina, et al.
Published: (2024)
by: Runkel, Christina, et al.
Published: (2024)
TANGO: Graph Neural Dynamics via Learned Energy and Tangential Flows
by: Eliasof, Moshe, et al.
Published: (2025)
by: Eliasof, Moshe, et al.
Published: (2025)
Approximation Theory for Lipschitz Continuous Transformers
by: Furuya, Takashi, et al.
Published: (2026)
by: Furuya, Takashi, et al.
Published: (2026)
Approximation theory for 1-Lipschitz ResNets
by: Murari, Davide, et al.
Published: (2025)
by: Murari, Davide, et al.
Published: (2025)
On the Effectiveness of Random Weights in Graph Neural Networks
by: Bui, Thu, et al.
Published: (2025)
by: Bui, Thu, et al.
Published: (2025)
laplax -- Laplace Approximations with JAX
by: Weber, Tobias, et al.
Published: (2025)
by: Weber, Tobias, et al.
Published: (2025)
Stochastic Multiresolution Image Sketching for Inverse Imaging Problems
by: Perelli, Alessandro, et al.
Published: (2024)
by: Perelli, Alessandro, et al.
Published: (2024)
Proximal Langevin Sampling With Inexact Proximal Mapping
by: Ehrhardt, Matthias J., et al.
Published: (2023)
by: Ehrhardt, Matthias J., et al.
Published: (2023)
Similar Items
-
Data-driven approaches to inverse problems
by: Schönlieb, Carola-Bibiane, et al.
Published: (2025) -
Rethinking Approximate Gaussian Inference in Classification
by: Mucsányi, Bálint, et al.
Published: (2025) -
Geometric Gaussian Approximations of Probability Distributions
by: Da Costa, Nathaël, et al.
Published: (2025) -
When is a System Discoverable from Data? Discovery Requires Chaos
by: Shumaylov, Zakhar, et al.
Published: (2025) -
Generalized Lie Symmetries in Physics-Informed Neural Operators
by: Wang, Amy Xiang, et al.
Published: (2025)