HarMoEny: Efficient Multi-GPU Inference of MoE Models
Fuente:
arXiv
Saved in:
| Main Authors: | Doucet, Zachary, Sharma, Rishi, de Vos, Martijn, Pires, Rafael, Kermarrec, Anne-Marie, Balmau, Oana |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Accelerating MoE Model Inference with Expert Sharding
by: Balmau, Oana, et al.
Published: (2025)
by: Balmau, Oana, et al.
Published: (2025)
Boosting Asynchronous Decentralized Learning with Model Fragmentation
by: Biswas, Sayan, et al.
Published: (2024)
by: Biswas, Sayan, et al.
Published: (2024)
Harnessing Increased Client Participation with Cohort-Parallel Federated Learning
by: Dhasade, Akash, et al.
Published: (2024)
by: Dhasade, Akash, et al.
Published: (2024)
Practical Federated Learning without a Server
by: Dhasade, Akash, et al.
Published: (2025)
by: Dhasade, Akash, et al.
Published: (2025)
Decentralized Learning Made Practical with Client Sampling
by: de Vos, Martijn, et al.
Published: (2023)
by: de Vos, Martijn, et al.
Published: (2023)
Fair Decentralized Learning
by: Biswas, Sayan, et al.
Published: (2024)
by: Biswas, Sayan, et al.
Published: (2024)
Energy-Aware Decentralized Learning with Intermittent Model Training
by: Dhasade, Akash, et al.
Published: (2024)
by: Dhasade, Akash, et al.
Published: (2024)
PeerSwap: A Peer-Sampler with Randomness Guarantees
by: Guerraoui, Rachid, et al.
Published: (2024)
by: Guerraoui, Rachid, et al.
Published: (2024)
Efficient CPU-GPU Collaborative Inference for MoE-based LLMs on Memory-Limited Systems
by: Huang, En-Ming, et al.
Published: (2025)
by: Huang, En-Ming, et al.
Published: (2025)
GRACE-MoE: Grouping and Replication with Locality-Aware Routing for Efficient Distributed MoE Inference
by: Han, Yu, et al.
Published: (2025)
by: Han, Yu, et al.
Published: (2025)
Decentralized Learning Made Easy with DecentralizePy
by: Dhasade, Akash, et al.
Published: (2023)
by: Dhasade, Akash, et al.
Published: (2023)
HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE Inference
by: Zhong, Shuzhang, et al.
Published: (2025)
by: Zhong, Shuzhang, et al.
Published: (2025)
MoE-Gen: High-Throughput MoE Inference on a Single GPU with Module-Based Batching
by: Xu, Tairan, et al.
Published: (2025)
by: Xu, Tairan, et al.
Published: (2025)
SP-MoE: Speculative Decoding and Prefetching for Accelerating MoE-based Model Inference
by: Chen, Liangkun, et al.
Published: (2025)
by: Chen, Liangkun, et al.
Published: (2025)
OD-MoE: On-Demand Expert Loading for Cacheless Edge-Distributed MoE Inference
by: Wang, Liujianfu, et al.
Published: (2025)
by: Wang, Liujianfu, et al.
Published: (2025)
Efficient Pyramidal Analysis of Gigapixel Images on a Decentralized Modest Computer Cluster
by: Reinbigler, Marie, et al.
Published: (2025)
by: Reinbigler, Marie, et al.
Published: (2025)
EPS-MoE: Expert Pipeline Scheduler for Cost-Efficient MoE Inference
by: Qian, Yulei, et al.
Published: (2024)
by: Qian, Yulei, et al.
Published: (2024)
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems
by: Wu, Qi, et al.
Published: (2026)
by: Wu, Qi, et al.
Published: (2026)
Get More for Less in Decentralized Learning Systems
by: Dhasade, Akash, et al.
Published: (2023)
by: Dhasade, Akash, et al.
Published: (2023)
Noiseless Privacy-Preserving Decentralized Learning
by: Biswas, Sayan, et al.
Published: (2024)
by: Biswas, Sayan, et al.
Published: (2024)
Accelerating Distributed MoE Training and Inference with Lina
by: Li, Jiamin, et al.
Published: (2022)
by: Li, Jiamin, et al.
Published: (2022)
Boosting Resource-Constrained Federated Learning Systems with Guessed Updates
by: Boukhari, Mohamed Yassine, et al.
Published: (2021)
by: Boukhari, Mohamed Yassine, et al.
Published: (2021)
MoE-Hub: Taming Software Complexity for Seamless MoE Overlap with Hardware-Accelerated Communication on Multi-GPU Systems
by: Zhou, Zhuoshan, et al.
Published: (2026)
by: Zhou, Zhuoshan, et al.
Published: (2026)
Staleness-Centric Optimizations for Parallel Diffusion MoE Inference
by: Luo, Jiajun, et al.
Published: (2024)
by: Luo, Jiajun, et al.
Published: (2024)
Janus: Disaggregating Attention and Experts for Scalable MoE Inference
by: Zhang, Zhexiang, et al.
Published: (2025)
by: Zhang, Zhexiang, et al.
Published: (2025)
Efficient Federated Search for Retrieval-Augmented Generation using Lightweight Routing
by: Dhasade, Akash, et al.
Published: (2025)
by: Dhasade, Akash, et al.
Published: (2025)
LSH-MoE: Communication-efficient MoE Training via Locality-Sensitive Hashing
by: Nie, Xiaonan, et al.
Published: (2024)
by: Nie, Xiaonan, et al.
Published: (2024)
ReviveMoE: Fast Recovery for Hardware Failures in Large-Scale MoE LLM Inference Deployments
by: Li, Haley, et al.
Published: (2026)
by: Li, Haley, et al.
Published: (2026)
MinatoLoader: Accelerating Machine Learning Training Through Efficient Data Preprocessing
by: Nouaji, Rahma, et al.
Published: (2025)
by: Nouaji, Rahma, et al.
Published: (2025)
Revisiting Ensembling in One-Shot Federated Learning
by: Allouah, Youssef, et al.
Published: (2024)
by: Allouah, Youssef, et al.
Published: (2024)
MoE-Compression: How the Compression Error of Experts Affects the Inference Accuracy of MoE Model?
by: Ma, Songkai, et al.
Published: (2025)
by: Ma, Songkai, et al.
Published: (2025)
MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism
by: Zhang, Zheng, et al.
Published: (2025)
by: Zhang, Zheng, et al.
Published: (2025)
Multi-Layer Scheduling for MoE-Based LLM Reasoning
by: Sun, Yifan, et al.
Published: (2026)
by: Sun, Yifan, et al.
Published: (2026)
ReaLB: Real-Time Load Balancing for Multimodal MoE Inference
by: Wang, Yingping, et al.
Published: (2026)
by: Wang, Yingping, et al.
Published: (2026)
Surviving Partial Rank Failures in Wide Expert-Parallel MoE Inference
by: Sun, Xun, et al.
Published: (2026)
by: Sun, Xun, et al.
Published: (2026)
A Pragmatic Approach to Learned Indexing in RocksDB: Targeted Optimizations with Minimal System Modification
by: Vashisth, Shubham, et al.
Published: (2026)
by: Vashisth, Shubham, et al.
Published: (2026)
Hexa-MoE: Efficient and Heterogeneous-aware Training for Mixture-of-Experts
by: Luo, Shuqing, et al.
Published: (2024)
by: Luo, Shuqing, et al.
Published: (2024)
Low-Cost Privacy-Preserving Decentralized Learning
by: Biswas, Sayan, et al.
Published: (2024)
by: Biswas, Sayan, et al.
Published: (2024)
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement
by: Wu, Tian, et al.
Published: (2025)
by: Wu, Tian, et al.
Published: (2025)
eMoE: Task-aware Memory Efficient Mixture-of-Experts-Based (MoE) Model Inference
by: Tairin, Suraiya, et al.
Published: (2025)
by: Tairin, Suraiya, et al.
Published: (2025)
Similar Items
-
Accelerating MoE Model Inference with Expert Sharding
by: Balmau, Oana, et al.
Published: (2025) -
Boosting Asynchronous Decentralized Learning with Model Fragmentation
by: Biswas, Sayan, et al.
Published: (2024) -
Harnessing Increased Client Participation with Cohort-Parallel Federated Learning
by: Dhasade, Akash, et al.
Published: (2024) -
Practical Federated Learning without a Server
by: Dhasade, Akash, et al.
Published: (2025) -
Decentralized Learning Made Practical with Client Sampling
by: de Vos, Martijn, et al.
Published: (2023)