Dion2: A Simple Method to Shrink Matrix in Muon
Fuente:
arXiv
Saved in:
| Main Authors: | Ahn, Kwangjun, Amsel, Noah, Langford, John |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SignMuon: Communication-Efficient Distributed Muon Optimization
by: Mishra, Neel, et al.
Published: (2026)
by: Mishra, Neel, et al.
Published: (2026)
Distributed Deep Learning using Stochastic Gradient Staleness
by: Pham, Viet Hoang, et al.
Published: (2025)
by: Pham, Viet Hoang, et al.
Published: (2025)
A Machine Learning Approach Towards Runtime Optimisation of Matrix Multiplication
by: Xia, Yufan, et al.
Published: (2026)
by: Xia, Yufan, et al.
Published: (2026)
Fast and Efficient 2-bit LLM Inference on GPU: 2/4/16-bit in a Weight Matrix with Asynchronous Dequantization
by: Li, Jinhao, et al.
Published: (2023)
by: Li, Jinhao, et al.
Published: (2023)
Canzona: A Unified, Asynchronous, and Load-Balanced Framework for Distributed Matrix-based Optimizers
by: Wang, Liangyu, et al.
Published: (2026)
by: Wang, Liangyu, et al.
Published: (2026)
From Tokens to Layers: Redefining Stall-Free Scheduling for MoE Serving with Layered Prefill
by: Lee, Gunjun, et al.
Published: (2025)
by: Lee, Gunjun, et al.
Published: (2025)
Cornserve: A Distributed Serving System for Any-to-Any Multimodal Models
by: Chung, Jae-Won, et al.
Published: (2026)
by: Chung, Jae-Won, et al.
Published: (2026)
LLMBridge: Reducing Costs to Access LLMs in a Prompt-Centric Internet
by: Martin, Noah, et al.
Published: (2024)
by: Martin, Noah, et al.
Published: (2024)
Towards providing reliable job completion time predictions using PCS
by: Faisal, Abdullah Bin, et al.
Published: (2024)
by: Faisal, Abdullah Bin, et al.
Published: (2024)
Spectral Sentinel: Scalable Byzantine-Robust Decentralized Federated Learning via Sketched Random Matrix Theory on Blockchain
by: Mishra, Animesh
Published: (2025)
by: Mishra, Animesh
Published: (2025)
A Comprehensive Survey of Federated Transfer Learning: Challenges, Methods and Applications
by: Guo, Wei, et al.
Published: (2024)
by: Guo, Wei, et al.
Published: (2024)
A Semantic Partitioning Method for Large-Scale Training of Knowledge Graph Embeddings
by: Bai, Yuhe
Published: (2025)
by: Bai, Yuhe
Published: (2025)
Nesterov Method for Asynchronous Pipeline Parallel Optimization
by: Ajanthan, Thalaiyasingam, et al.
Published: (2025)
by: Ajanthan, Thalaiyasingam, et al.
Published: (2025)
MIRA: A Method of Federated MultI-Task Learning for LaRge LAnguage Models
by: Elbakary, Ahmed, et al.
Published: (2024)
by: Elbakary, Ahmed, et al.
Published: (2024)
Cornfigurator: Automated Planning for Any-to-Any Multimodal Model Serving
by: Ma, Jeff J., et al.
Published: (2025)
by: Ma, Jeff J., et al.
Published: (2025)
Unlearning during Learning: An Efficient Federated Machine Unlearning Method
by: Gu, Hanlin, et al.
Published: (2024)
by: Gu, Hanlin, et al.
Published: (2024)
A Light-weight and Unsupervised Method for Near Real-time Behavioral Analysis using Operational Data Measurement
by: Vargis, Tom Richard, et al.
Published: (2024)
by: Vargis, Tom Richard, et al.
Published: (2024)
Robust Fully-Asynchronous Methods for Distributed Training over General Architecture
by: Zhu, Zehan, et al.
Published: (2023)
by: Zhu, Zehan, et al.
Published: (2023)
Leiden-Fusion Partitioning Method for Effective Distributed Training of Graph Embeddings
by: Bai, Yuhe, et al.
Published: (2024)
by: Bai, Yuhe, et al.
Published: (2024)
Unlocking Dynamic Inter-Client Spatial Dependencies: A Federated Spatio-Temporal Graph Learning Method for Traffic Flow Forecasting
by: Wang, Feng, et al.
Published: (2025)
by: Wang, Feng, et al.
Published: (2025)
FedUHB: Accelerating Federated Unlearning via Polyak Heavy Ball Method
by: Jiang, Yu, et al.
Published: (2024)
by: Jiang, Yu, et al.
Published: (2024)
Distributed Matrix-Based Sampling for Graph Neural Network Training
by: Tripathy, Alok, et al.
Published: (2023)
by: Tripathy, Alok, et al.
Published: (2023)
First-Order Softmax Weighted Switching Gradient Method for Distributed Stochastic Minimax Optimization with Stochastic Constraints
by: Luo, Zhankun, et al.
Published: (2026)
by: Luo, Zhankun, et al.
Published: (2026)
Simple Opinion Dynamics for No-Regret Learning
by: Lazarsfeld, John, et al.
Published: (2023)
by: Lazarsfeld, John, et al.
Published: (2023)
Fast Matrix Multiplications for Lookup Table-Quantized LLMs
by: Guo, Han, et al.
Published: (2024)
by: Guo, Han, et al.
Published: (2024)
INTELLECT-2: A Reasoning Model Trained Through Globally Decentralized Reinforcement Learning
by: Prime Intellect Team, et al.
Published: (2025)
by: Prime Intellect Team, et al.
Published: (2025)
ECHO-2: A Large-Scale Distributed Rollout Framework for Cost-Efficient Reinforcement Learning
by: Song, Jingwei, et al.
Published: (2026)
by: Song, Jingwei, et al.
Published: (2026)
MAS-H2: A Hierarchical Multi-Agent System for Holistic Cloud-Native Autoscaling
by: Hamzeh, Hamed, et al.
Published: (2026)
by: Hamzeh, Hamed, et al.
Published: (2026)
OpenG2G: A Simulation Platform for AI Datacenter-Grid Runtime Coordination
by: Chung, Jae-Won, et al.
Published: (2026)
by: Chung, Jae-Won, et al.
Published: (2026)
DiffKV: Differentiated Memory Management for Large Language Models with Parallel KV Compaction
by: Zhang, Yanqi, et al.
Published: (2024)
by: Zhang, Yanqi, et al.
Published: (2024)
Accelerating Sparse Matrix-Matrix Multiplication on GPUs with Processing Near HBMs
by: Li, Shiju, et al.
Published: (2025)
by: Li, Shiju, et al.
Published: (2025)
Secure Federated Learning Across Heterogeneous Cloud and High-Performance Computing Resources -- A Case Study on Federated Fine-tuning of LLaMA 2
by: Li, Zilinghan, et al.
Published: (2024)
by: Li, Zilinghan, et al.
Published: (2024)
Asyn2F: An Asynchronous Federated Learning Framework with Bidirectional Model Aggregation
by: Cao, Tien-Dung, et al.
Published: (2024)
by: Cao, Tien-Dung, et al.
Published: (2024)
DHO$_2$: Accelerating Distributed Hybrid Order Optimization via Model Parallelism and ADMM
by: Gu, Shunxian, et al.
Published: (2025)
by: Gu, Shunxian, et al.
Published: (2025)
ASTRA-sim2.0: Modeling Hierarchical Networks and Disaggregated Systems for Large-model Training at Scale
by: Won, William, et al.
Published: (2023)
by: Won, William, et al.
Published: (2023)
Revisiting Reliability in Large-Scale Machine Learning Research Clusters
by: Kokolis, Apostolos, et al.
Published: (2024)
by: Kokolis, Apostolos, et al.
Published: (2024)
A Bayesian Framework for Clustered Federated Learning
by: Wu, Peng, et al.
Published: (2024)
by: Wu, Peng, et al.
Published: (2024)
A Theory of Multi-Agent Generative Flow Networks
by: Brunswic, Leo Maxime, et al.
Published: (2025)
by: Brunswic, Leo Maxime, et al.
Published: (2025)
Drift-Aware Federated Learning: A Causal Perspective
by: Fang, Yunjie, et al.
Published: (2025)
by: Fang, Yunjie, et al.
Published: (2025)
A Survey on Contribution Evaluation in Vertical Federated Learning
by: Cui, Yue, et al.
Published: (2024)
by: Cui, Yue, et al.
Published: (2024)
Similar Items
-
SignMuon: Communication-Efficient Distributed Muon Optimization
by: Mishra, Neel, et al.
Published: (2026) -
Distributed Deep Learning using Stochastic Gradient Staleness
by: Pham, Viet Hoang, et al.
Published: (2025) -
A Machine Learning Approach Towards Runtime Optimisation of Matrix Multiplication
by: Xia, Yufan, et al.
Published: (2026) -
Fast and Efficient 2-bit LLM Inference on GPU: 2/4/16-bit in a Weight Matrix with Asynchronous Dequantization
by: Li, Jinhao, et al.
Published: (2023) -
Canzona: A Unified, Asynchronous, and Load-Balanced Framework for Distributed Matrix-based Optimizers
by: Wang, Liangyu, et al.
Published: (2026)