Optimal Transport Aggregation for Distributed Mixture-of-Experts
Fuente:
arXiv
Saved in:
| Main Authors: | Chamroukhi, Faïcel, Pham, Nhat Thien |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Score Distributions to Balance: Plug-and-Play Mixture-of-Experts Routing
by: Shahout, Rana, et al.
Published: (2025)
by: Shahout, Rana, et al.
Published: (2025)
Mosaic: Data-Free Knowledge Distillation via Mixture-of-Experts for Heterogeneous Distributed Environments
by: Liu, Junming, et al.
Published: (2025)
by: Liu, Junming, et al.
Published: (2025)
Utility-Driven Speculative Decoding for Mixture-of-Experts
by: Saxena, Anish, et al.
Published: (2025)
by: Saxena, Anish, et al.
Published: (2025)
Exploiting Inter-Layer Expert Affinity for Accelerating Mixture-of-Experts Model Inference
by: Yao, Jinghan, et al.
Published: (2024)
by: Yao, Jinghan, et al.
Published: (2024)
A Survey on Inference Optimization Techniques for Mixture of Experts Models
by: Liu, Jiacheng, et al.
Published: (2024)
by: Liu, Jiacheng, et al.
Published: (2024)
Comet: Fine-grained Computation-communication Overlapping for Mixture-of-Experts
by: Zhang, Shulai, et al.
Published: (2025)
by: Zhang, Shulai, et al.
Published: (2025)
RaMP: Runtime-Aware Megakernel Polymorphism for Mixture-of-Experts
by: Sharma, Vyom, et al.
Published: (2026)
by: Sharma, Vyom, et al.
Published: (2026)
Mixture of Experts with Mixture of Precisions for Tuning Quality of Service
by: Imani, HamidReza, et al.
Published: (2024)
by: Imani, HamidReza, et al.
Published: (2024)
Scalable Pretraining of Large Mixture of Experts Language Models on Aurora Super Computer
by: Vooturi, Dharma Teja, et al.
Published: (2026)
by: Vooturi, Dharma Teja, et al.
Published: (2026)
Adaptive Consensus Gradients Aggregation for Scaled Distributed Training
by: Choukroun, Yoni, et al.
Published: (2024)
by: Choukroun, Yoni, et al.
Published: (2024)
WDMoE: Wireless Distributed Mixture of Experts for Large Language Models
by: Xue, Nan, et al.
Published: (2024)
by: Xue, Nan, et al.
Published: (2024)
Global and Local Prompts Cooperation via Optimal Transport for Federated Learning
by: Li, Hongxia, et al.
Published: (2024)
by: Li, Hongxia, et al.
Published: (2024)
BlackMamba: Mixture of Experts for State-Space Models
by: Anthony, Quentin, et al.
Published: (2024)
by: Anthony, Quentin, et al.
Published: (2024)
Linear-MoE: Linear Sequence Modeling Meets Mixture-of-Experts
by: Sun, Weigao, et al.
Published: (2025)
by: Sun, Weigao, et al.
Published: (2025)
FourierMoE: Fourier Mixture-of-Experts Adaptation of Large Language Models
by: Jiang, Juyong, et al.
Published: (2026)
by: Jiang, Juyong, et al.
Published: (2026)
OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models
by: Xue, Fuzhao, et al.
Published: (2024)
by: Xue, Fuzhao, et al.
Published: (2024)
HybridEP: Scaling Expert Parallelism to Cross-Datacenter Scenario via Hybrid Expert/Data Transmission
by: Yang, Weihao, et al.
Published: (2025)
by: Yang, Weihao, et al.
Published: (2025)
FedAH: Aggregated Head for Personalized Federated Learning
by: Zhou, Pengzhan, et al.
Published: (2024)
by: Zhou, Pengzhan, et al.
Published: (2024)
Beyond Aggregation: Guiding Clients in Heterogeneous Federated Learning
by: Wang, Zijian, et al.
Published: (2025)
by: Wang, Zijian, et al.
Published: (2025)
Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models
by: Kamahori, Keisuke, et al.
Published: (2024)
by: Kamahori, Keisuke, et al.
Published: (2024)
Accelerating MoE Model Inference with Expert Sharding
by: Balmau, Oana, et al.
Published: (2025)
by: Balmau, Oana, et al.
Published: (2025)
Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing
by: Liu, Mengfan, et al.
Published: (2025)
by: Liu, Mengfan, et al.
Published: (2025)
FedMoE-DA: Federated Mixture of Experts via Domain Aware Fine-grained Aggregation
by: Zhan, Ziwei, et al.
Published: (2024)
by: Zhan, Ziwei, et al.
Published: (2024)
EASTER: Embedding Aggregation-based Heterogeneous Models Training in Vertical Federated Learning
by: Wang, Shuo, et al.
Published: (2023)
by: Wang, Shuo, et al.
Published: (2023)
MoSKA: Mixture of Shared KV Attention for Efficient Long-Sequence LLM Inference
by: Rhee, Myunghyun, et al.
Published: (2025)
by: Rhee, Myunghyun, et al.
Published: (2025)
MoESys: A Distributed and Efficient Mixture-of-Experts Training and Inference System for Internet Services
by: Yu, Dianhai, et al.
Published: (2022)
by: Yu, Dianhai, et al.
Published: (2022)
SEAFL: Enhancing Efficiency in Semi-Asynchronous Federated Learning through Adaptive Aggregation and Selective Training
by: Islam, Md Sirajul, et al.
Published: (2025)
by: Islam, Md Sirajul, et al.
Published: (2025)
FedAA: A Reinforcement Learning Perspective on Adaptive Aggregation for Fair and Robust Federated Learning
by: He, Jialuo, et al.
Published: (2024)
by: He, Jialuo, et al.
Published: (2024)
Scattered Mixture-of-Experts Implementation
by: Tan, Shawn, et al.
Published: (2024)
by: Tan, Shawn, et al.
Published: (2024)
AgentStop: Terminating Local AI Agents Early to Save Energy in Consumer Devices
by: Pham, Dzung, et al.
Published: (2026)
by: Pham, Dzung, et al.
Published: (2026)
Taming Latency-Memory Trade-Off in MoE-Based LLM Serving via Fine-Grained Expert Offloading
by: Yu, Hanfei, et al.
Published: (2025)
by: Yu, Hanfei, et al.
Published: (2025)
Federated Hierarchical Clustering with Automatic Selection of Optimal Cluster Numbers
by: Zhang, Yue, et al.
Published: (2026)
by: Zhang, Yue, et al.
Published: (2026)
Trustworthiness of Stochastic Gradient Descent in Distributed Learning
by: Li, Hongyang, et al.
Published: (2024)
by: Li, Hongyang, et al.
Published: (2024)
On the Fragility of Data Attribution When Learning Is Distributed
by: Gao, Xian, et al.
Published: (2026)
by: Gao, Xian, et al.
Published: (2026)
Loss- and Reward-Weighting for Efficient Distributed Reinforcement Learning
by: Holen, Martin, et al.
Published: (2023)
by: Holen, Martin, et al.
Published: (2023)
Measuring Heterogeneity in Machine Learning with Distributed Energy Distance
by: Fan, Mengchen, et al.
Published: (2025)
by: Fan, Mengchen, et al.
Published: (2025)
Distributed Low-Communication Training with Decoupled Momentum Optimization
by: Nedelkoski, Sasho, et al.
Published: (2025)
by: Nedelkoski, Sasho, et al.
Published: (2025)
PALS: Power-Aware LLM Serving for Mixture-of-Experts Models
by: Hankendi, Can, et al.
Published: (2026)
by: Hankendi, Can, et al.
Published: (2026)
TrainVerify: Equivalence-Based Verification for Distributed LLM Training
by: Lu, Yunchi, et al.
Published: (2025)
by: Lu, Yunchi, et al.
Published: (2025)
Galvatron: An Automatic Distributed System for Efficient Foundation Model Training
by: Liu, Xinyi, et al.
Published: (2025)
by: Liu, Xinyi, et al.
Published: (2025)
Similar Items
-
From Score Distributions to Balance: Plug-and-Play Mixture-of-Experts Routing
by: Shahout, Rana, et al.
Published: (2025) -
Mosaic: Data-Free Knowledge Distillation via Mixture-of-Experts for Heterogeneous Distributed Environments
by: Liu, Junming, et al.
Published: (2025) -
Utility-Driven Speculative Decoding for Mixture-of-Experts
by: Saxena, Anish, et al.
Published: (2025) -
Exploiting Inter-Layer Expert Affinity for Accelerating Mixture-of-Experts Model Inference
by: Yao, Jinghan, et al.
Published: (2024) -
A Survey on Inference Optimization Techniques for Mixture of Experts Models
by: Liu, Jiacheng, et al.
Published: (2024)