Gespeichert in:
| Hauptverfasser: | Massena, Thomas, Friedrich, Corentin, Serrurier, Mathieu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2605.19781 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Turbo-Muon: Accelerating Orthogonality-Based Optimization with Pre-Conditioning
von: Boissin, Thibaut, et al.
Veröffentlicht: (2025)
von: Boissin, Thibaut, et al.
Veröffentlicht: (2025)
Efficient Robust Conformal Prediction via Lipschitz-Bounded Networks
von: Massena, Thomas, et al.
Veröffentlicht: (2025)
von: Massena, Thomas, et al.
Veröffentlicht: (2025)
An Adaptive Orthogonal Convolution Scheme for Efficient and Flexible CNN Architectures
von: Boissin, Thibaut, et al.
Veröffentlicht: (2025)
von: Boissin, Thibaut, et al.
Veröffentlicht: (2025)
Fast and Flexible Robustness Certificates for Semantic Segmentation
von: Massena, Thomas, et al.
Veröffentlicht: (2025)
von: Massena, Thomas, et al.
Veröffentlicht: (2025)
DP-SGD Without Clipping: The Lipschitz Neural Network Way
von: Bethune, Louis, et al.
Veröffentlicht: (2023)
von: Bethune, Louis, et al.
Veröffentlicht: (2023)
Generating Heterogeneous Multi-dimensional Data : A Comparative Study
von: Corbeau, Michael, et al.
Veröffentlicht: (2025)
von: Corbeau, Michael, et al.
Veröffentlicht: (2025)
On the explainable properties of 1-Lipschitz Neural Networks: An Optimal Transport Perspective
von: Serrurier, Mathieu, et al.
Veröffentlicht: (2022)
von: Serrurier, Mathieu, et al.
Veröffentlicht: (2022)
DC-SGD: Differentially Private SGD with Dynamic Clipping through Gradient Norm Distribution Estimation
von: Wei, Chengkun, et al.
Veröffentlicht: (2025)
von: Wei, Chengkun, et al.
Veröffentlicht: (2025)
Muon$^2$: Boosting Muon via Adaptive Second-Moment Preconditioning
von: Liu, Ziyue, et al.
Veröffentlicht: (2026)
von: Liu, Ziyue, et al.
Veröffentlicht: (2026)
TrasMuon: Trust-Region Adaptive Scaling for Orthogonalized Momentum Optimizers
von: Cheng, Peng, et al.
Veröffentlicht: (2026)
von: Cheng, Peng, et al.
Veröffentlicht: (2026)
PowerStep: Memory-Efficient Adaptive Optimization via $\ell_p$-Norm Steepest Descent
von: Lu, Yao, et al.
Veröffentlicht: (2026)
von: Lu, Yao, et al.
Veröffentlicht: (2026)
Anon: Extrapolating Adaptivity Beyond SGD and Adam
von: Zhang, Yiheng, et al.
Veröffentlicht: (2026)
von: Zhang, Yiheng, et al.
Veröffentlicht: (2026)
The Newton-Muon Optimizer
von: Du, Zhehang, et al.
Veröffentlicht: (2026)
von: Du, Zhehang, et al.
Veröffentlicht: (2026)
Optimizer-Induced Mode Connectivity: From AdamW to Muon
von: Zhang, Fangzhao, et al.
Veröffentlicht: (2026)
von: Zhang, Fangzhao, et al.
Veröffentlicht: (2026)
MiMuon: Mixed Muon Optimizer with Improved Generalization for Large Models
von: Huang, Feihu, et al.
Veröffentlicht: (2026)
von: Huang, Feihu, et al.
Veröffentlicht: (2026)
Muon Optimizer Accelerates Grokking
von: Tveit, Amund, et al.
Veröffentlicht: (2025)
von: Tveit, Amund, et al.
Veröffentlicht: (2025)
Diagonalisation SGD: Fast & Convergent SGD for Non-Differentiable Models via Reparameterisation and Smoothing
von: Wagner, Dominik, et al.
Veröffentlicht: (2024)
von: Wagner, Dominik, et al.
Veröffentlicht: (2024)
Intrinsic Muon: Spectral Optimization on Riemannian Matrix Manifolds
von: Li, Yibang, et al.
Veröffentlicht: (2026)
von: Li, Yibang, et al.
Veröffentlicht: (2026)
Enhancing DP-SGD through Non-monotonous Adaptive Scaling Gradient Weight
von: Huang, Tao, et al.
Veröffentlicht: (2024)
von: Huang, Tao, et al.
Veröffentlicht: (2024)
Comparative Analysis of Novel NIRMAL Optimizer Against Adam and SGD with Momentum
von: Gaud, Nirmal, et al.
Veröffentlicht: (2025)
von: Gaud, Nirmal, et al.
Veröffentlicht: (2025)
DeMuon: A Decentralized Muon for Matrix Optimization over Graphs
von: He, Chuan, et al.
Veröffentlicht: (2025)
von: He, Chuan, et al.
Veröffentlicht: (2025)
The Ky Fan Norms and Beyond: Dual Norms and Combinations for Matrix Optimization
von: Kravatskiy, Alexey, et al.
Veröffentlicht: (2025)
von: Kravatskiy, Alexey, et al.
Veröffentlicht: (2025)
DynMuon: A Dynamic Spectral Shaping View of Muon
von: Wu, Fangzhou, et al.
Veröffentlicht: (2026)
von: Wu, Fangzhou, et al.
Veröffentlicht: (2026)
Connections between Schedule-Free Optimizers, AdEMAMix, and Accelerated SGD Variants
von: Morwani, Depen, et al.
Veröffentlicht: (2025)
von: Morwani, Depen, et al.
Veröffentlicht: (2025)
Representations learnt by SGD and Adaptive learning rules: Conditions that vary sparsity and selectivity in neural networks
von: Park, Jin Hyun
Veröffentlicht: (2022)
von: Park, Jin Hyun
Veröffentlicht: (2022)
Advancing Model Refinement: Muon-Optimized Distillation and Quantization for LLM Deployment
von: Sander, Jacob, et al.
Veröffentlicht: (2026)
von: Sander, Jacob, et al.
Veröffentlicht: (2026)
Scaling Laws of SignSGD in Linear Regression: When Does It Outperform SGD?
von: Kim, Jihwan, et al.
Veröffentlicht: (2026)
von: Kim, Jihwan, et al.
Veröffentlicht: (2026)
Bootstrap SGD: Algorithmic Stability and Robustness
von: Christmann, Andreas, et al.
Veröffentlicht: (2024)
von: Christmann, Andreas, et al.
Veröffentlicht: (2024)
HTMuon: Improving Muon via Heavy-Tailed Spectral Correction
von: Pang, Tianyu, et al.
Veröffentlicht: (2026)
von: Pang, Tianyu, et al.
Veröffentlicht: (2026)
RQP-SGD: Differential Private Machine Learning through Noisy SGD and Randomized Quantization
von: Feng, Ce, et al.
Veröffentlicht: (2024)
von: Feng, Ce, et al.
Veröffentlicht: (2024)
MuonRec: Shifting the Optimizer Paradigm Beyond Adam in Scalable Generative Recommendation
von: Shan, Rong, et al.
Veröffentlicht: (2026)
von: Shan, Rong, et al.
Veröffentlicht: (2026)
Adaptive Accountability in Networked MAS: Tracing and Mitigating Emergent Norms at Scale
von: Alqithami, Saad
Veröffentlicht: (2025)
von: Alqithami, Saad
Veröffentlicht: (2025)
Exploring Sparsity and Smoothness of Arbitrary $\ell_p$ Norms in Adversarial Attacks
von: Duhme, Christof, et al.
Veröffentlicht: (2026)
von: Duhme, Christof, et al.
Veröffentlicht: (2026)
Accumulative SGD Influence Estimation for Data Attribution
von: Shi, Yunxiao, et al.
Veröffentlicht: (2025)
von: Shi, Yunxiao, et al.
Veröffentlicht: (2025)
Fuzzy Representation of Norms
von: Assadi, Ziba, et al.
Veröffentlicht: (2026)
von: Assadi, Ziba, et al.
Veröffentlicht: (2026)
StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models
von: Yu, Dingzhi, et al.
Veröffentlicht: (2026)
von: Yu, Dingzhi, et al.
Veröffentlicht: (2026)
Decoupled Weight Decay for Any $p$ Norm
von: Outmezguine, Nadav Joseph, et al.
Veröffentlicht: (2024)
von: Outmezguine, Nadav Joseph, et al.
Veröffentlicht: (2024)
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam
von: Peng, Hanyang, et al.
Veröffentlicht: (2025)
von: Peng, Hanyang, et al.
Veröffentlicht: (2025)
Worker Disagreement Reveals Sharp Directions in Local SGD
von: Dimlioglu, Tolga, et al.
Veröffentlicht: (2026)
von: Dimlioglu, Tolga, et al.
Veröffentlicht: (2026)
Muon is Scalable for LLM Training
von: Liu, Jingyuan, et al.
Veröffentlicht: (2025)
von: Liu, Jingyuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Turbo-Muon: Accelerating Orthogonality-Based Optimization with Pre-Conditioning
von: Boissin, Thibaut, et al.
Veröffentlicht: (2025) -
Efficient Robust Conformal Prediction via Lipschitz-Bounded Networks
von: Massena, Thomas, et al.
Veröffentlicht: (2025) -
An Adaptive Orthogonal Convolution Scheme for Efficient and Flexible CNN Architectures
von: Boissin, Thibaut, et al.
Veröffentlicht: (2025) -
Fast and Flexible Robustness Certificates for Semantic Segmentation
von: Massena, Thomas, et al.
Veröffentlicht: (2025) -
DP-SGD Without Clipping: The Lipschitz Neural Network Way
von: Bethune, Louis, et al.
Veröffentlicht: (2023)