Turbo-Muon: Accelerating Orthogonality-Based Optimization with Pre-Conditioning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Boissin, Thibaut, Massena, Thomas, Mamalet, Franck, Serrurier, Mathieu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
An Adaptive Orthogonal Convolution Scheme for Efficient and Flexible CNN Architectures
von: Boissin, Thibaut, et al.
Veröffentlicht: (2025)
von: Boissin, Thibaut, et al.
Veröffentlicht: (2025)
Efficient Robust Conformal Prediction via Lipschitz-Bounded Networks
von: Massena, Thomas, et al.
Veröffentlicht: (2025)
von: Massena, Thomas, et al.
Veröffentlicht: (2025)
On the explainable properties of 1-Lipschitz Neural Networks: An Optimal Transport Perspective
von: Serrurier, Mathieu, et al.
Veröffentlicht: (2022)
von: Serrurier, Mathieu, et al.
Veröffentlicht: (2022)
From SGD to Muon: Adaptive Optimization via Schatten-p Norms
von: Massena, Thomas, et al.
Veröffentlicht: (2026)
von: Massena, Thomas, et al.
Veröffentlicht: (2026)
Orthogonium : A Unified, Efficient Library of Orthogonal and 1-Lipschitz Building Blocks
von: Boissin, Thibaut, et al.
Veröffentlicht: (2026)
von: Boissin, Thibaut, et al.
Veröffentlicht: (2026)
Fast and Flexible Robustness Certificates for Semantic Segmentation
von: Massena, Thomas, et al.
Veröffentlicht: (2025)
von: Massena, Thomas, et al.
Veröffentlicht: (2025)
DP-SGD Without Clipping: The Lipschitz Neural Network Way
von: Bethune, Louis, et al.
Veröffentlicht: (2023)
von: Bethune, Louis, et al.
Veröffentlicht: (2023)
HadamRNN: Binary and Sparse Ternary Orthogonal RNNs
von: Foucault, Armand, et al.
Veröffentlicht: (2025)
von: Foucault, Armand, et al.
Veröffentlicht: (2025)
Quantized Approximately Orthogonal Recurrent Neural Networks
von: Foucault, Armand, et al.
Veröffentlicht: (2024)
von: Foucault, Armand, et al.
Veröffentlicht: (2024)
FedMuon: Accelerating Federated Learning with Matrix Orthogonalization
von: Liu, Junkang, et al.
Veröffentlicht: (2025)
von: Liu, Junkang, et al.
Veröffentlicht: (2025)
Generating Heterogeneous Multi-dimensional Data : A Comparative Study
von: Corbeau, Michael, et al.
Veröffentlicht: (2025)
von: Corbeau, Michael, et al.
Veröffentlicht: (2025)
Muon Optimizer Accelerates Grokking
von: Tveit, Amund, et al.
Veröffentlicht: (2025)
von: Tveit, Amund, et al.
Veröffentlicht: (2025)
TrasMuon: Trust-Region Adaptive Scaling for Orthogonalized Momentum Optimizers
von: Cheng, Peng, et al.
Veröffentlicht: (2026)
von: Cheng, Peng, et al.
Veröffentlicht: (2026)
Back to the Baseline: Examining Baseline Effects on Explainability Metrics
von: Picard, Agustin Martin, et al.
Veröffentlicht: (2025)
von: Picard, Agustin Martin, et al.
Veröffentlicht: (2025)
Preconditioning Benefits of Spectral Orthogonalization in Muon
von: Ma, Jianhao, et al.
Veröffentlicht: (2026)
von: Ma, Jianhao, et al.
Veröffentlicht: (2026)
Turbo-ICL: In-Context Learning-Based Turbo Equalization
von: Song, Zihang, et al.
Veröffentlicht: (2025)
von: Song, Zihang, et al.
Veröffentlicht: (2025)
TurboHopp: Accelerated Molecule Scaffold Hopping with Consistency Models
von: Yoo, Kiwoong, et al.
Veröffentlicht: (2024)
von: Yoo, Kiwoong, et al.
Veröffentlicht: (2024)
Unlocking Feature Visualization for Deeper Networks with MAgnitude Constrained Optimization
von: Fel, Thomas, et al.
Veröffentlicht: (2023)
von: Fel, Thomas, et al.
Veröffentlicht: (2023)
Spectral Flattening Is All Muon Needs: How Orthogonalization Controls Learning Rate and Convergence
von: Nguyen, Tien-Phat, et al.
Veröffentlicht: (2026)
von: Nguyen, Tien-Phat, et al.
Veröffentlicht: (2026)
The Newton-Muon Optimizer
von: Du, Zhehang, et al.
Veröffentlicht: (2026)
von: Du, Zhehang, et al.
Veröffentlicht: (2026)
MiMuon: Mixed Muon Optimizer with Improved Generalization for Large Models
von: Huang, Feihu, et al.
Veröffentlicht: (2026)
von: Huang, Feihu, et al.
Veröffentlicht: (2026)
Robust One-Class Classification with Signed Distance Function using 1-Lipschitz Neural Networks
von: Bethune, Louis, et al.
Veröffentlicht: (2023)
von: Bethune, Louis, et al.
Veröffentlicht: (2023)
Orthogonalized Policy Optimization:Policy Optimization as Orthogonal Projection in Hilbert Space
von: Zixian, Wang
Veröffentlicht: (2026)
von: Zixian, Wang
Veröffentlicht: (2026)
Turbo4DGen: Ultra-Fast Acceleration for 4D Generation
von: Man, Yuanbin, et al.
Veröffentlicht: (2026)
von: Man, Yuanbin, et al.
Veröffentlicht: (2026)
TurboSpec: Closed-loop Speculation Control System for Optimizing LLM Serving Goodput
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2024)
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2024)
TurboDiffusion: Accelerating Video Diffusion Models by 100-200 Times
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
Intrinsic Muon: Spectral Optimization on Riemannian Matrix Manifolds
von: Li, Yibang, et al.
Veröffentlicht: (2026)
von: Li, Yibang, et al.
Veröffentlicht: (2026)
Group Orthogonalized Policy Optimization:Group Policy Optimization as Orthogonal Projection in Hilbert Space
von: Zixian, Wang
Veröffentlicht: (2026)
von: Zixian, Wang
Veröffentlicht: (2026)
How to design a dataset compliant with an ML-based system ODD?
von: Cappi, Cyril, et al.
Veröffentlicht: (2024)
von: Cappi, Cyril, et al.
Veröffentlicht: (2024)
TEON: Tensorized Orthonormalization Beyond Layer-Wise Muon for Large Language Model Pre-Training
von: Zhang, Ruijie, et al.
Veröffentlicht: (2026)
von: Zhang, Ruijie, et al.
Veröffentlicht: (2026)
DynMuon: A Dynamic Spectral Shaping View of Muon
von: Wu, Fangzhou, et al.
Veröffentlicht: (2026)
von: Wu, Fangzhou, et al.
Veröffentlicht: (2026)
LARD 2.0: Enhanced Datasets and Benchmarking for Autonomous Landing Systems
von: Bougacha, Yassine, et al.
Veröffentlicht: (2026)
von: Bougacha, Yassine, et al.
Veröffentlicht: (2026)
ORTHOBO: Orthogonal Bayesian Hyperparameter Optimization
von: Schröder, Maresa, et al.
Veröffentlicht: (2026)
von: Schröder, Maresa, et al.
Veröffentlicht: (2026)
DeMuon: A Decentralized Muon for Matrix Optimization over Graphs
von: He, Chuan, et al.
Veröffentlicht: (2025)
von: He, Chuan, et al.
Veröffentlicht: (2025)
Advancing Model Refinement: Muon-Optimized Distillation and Quantization for LLM Deployment
von: Sander, Jacob, et al.
Veröffentlicht: (2026)
von: Sander, Jacob, et al.
Veröffentlicht: (2026)
Muon$^2$: Boosting Muon via Adaptive Second-Moment Preconditioning
von: Liu, Ziyue, et al.
Veröffentlicht: (2026)
von: Liu, Ziyue, et al.
Veröffentlicht: (2026)
AdaFuse: Accelerating Dynamic Adapter Inference via Token-Level Pre-Gating and Fused Kernel Optimization
von: Li, Qiyang, et al.
Veröffentlicht: (2026)
von: Li, Qiyang, et al.
Veröffentlicht: (2026)
TurboSAT: Gradient-Guided Boolean Satisfiability Accelerated on GPU-CPU Hybrid System
von: Dai, Steve, et al.
Veröffentlicht: (2025)
von: Dai, Steve, et al.
Veröffentlicht: (2025)
MuonRec: Shifting the Optimizer Paradigm Beyond Adam in Scalable Generative Recommendation
von: Shan, Rong, et al.
Veröffentlicht: (2026)
von: Shan, Rong, et al.
Veröffentlicht: (2026)
Refine and Purify: Orthogonal Basis Optimization with Null-Space Denoising for Conditional Representation Learning
von: Wang, Jiaquan, et al.
Veröffentlicht: (2026)
von: Wang, Jiaquan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
An Adaptive Orthogonal Convolution Scheme for Efficient and Flexible CNN Architectures
von: Boissin, Thibaut, et al.
Veröffentlicht: (2025) -
Efficient Robust Conformal Prediction via Lipschitz-Bounded Networks
von: Massena, Thomas, et al.
Veröffentlicht: (2025) -
On the explainable properties of 1-Lipschitz Neural Networks: An Optimal Transport Perspective
von: Serrurier, Mathieu, et al.
Veröffentlicht: (2022) -
From SGD to Muon: Adaptive Optimization via Schatten-p Norms
von: Massena, Thomas, et al.
Veröffentlicht: (2026) -
Orthogonium : A Unified, Efficient Library of Orthogonal and 1-Lipschitz Building Blocks
von: Boissin, Thibaut, et al.
Veröffentlicht: (2026)