BigMac: A Communication-Efficient Mixture-of-Experts Model Structure for Fast Training and Inference
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jin, Zewen, Wang, Shengnan, Zhu, Jiaan, Zhan, Hongrui, Bai, Youhui, Zhang, Lin, Ming, Zhenyu, Li, Cheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BigMac: Breaking the Pareto Frontier of Compute and Memory in Multimodal LLM Training
von: Zhang, Zili, et al.
Veröffentlicht: (2026)
von: Zhang, Zili, et al.
Veröffentlicht: (2026)
Accelerating Long-Tail Generation in Synchronous RLHF Training via Adaptive Tensor Parallelism
von: Zhao, Long, et al.
Veröffentlicht: (2026)
von: Zhao, Long, et al.
Veröffentlicht: (2026)
LiteCache: A Query Similarity-Driven, GPU-Centric KVCache Subsystem for Efficient LLM Inference
von: Yi, Jiawei, et al.
Veröffentlicht: (2025)
von: Yi, Jiawei, et al.
Veröffentlicht: (2025)
AdaCluster: Adaptive Query-Key Clustering for Sparse Attention in Video Generation
von: Tan, Haoyue, et al.
Veröffentlicht: (2026)
von: Tan, Haoyue, et al.
Veröffentlicht: (2026)
MoBiLE: Efficient Mixture-of-Experts Inference on Consumer GPU with Mixture of Big Little Experts
von: Zhao, Yushu, et al.
Veröffentlicht: (2025)
von: Zhao, Yushu, et al.
Veröffentlicht: (2025)
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference
von: Gong, Ping, et al.
Veröffentlicht: (2025)
von: Gong, Ping, et al.
Veröffentlicht: (2025)
MegaScale-MoE: Large-Scale Communication-Efficient Training of Mixture-of-Experts Models in Production
von: Jin, Chao, et al.
Veröffentlicht: (2025)
von: Jin, Chao, et al.
Veröffentlicht: (2025)
Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models
von: Kamahori, Keisuke, et al.
Veröffentlicht: (2024)
von: Kamahori, Keisuke, et al.
Veröffentlicht: (2024)
HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference
von: Lin, Haoran, et al.
Veröffentlicht: (2025)
von: Lin, Haoran, et al.
Veröffentlicht: (2025)
CARL-MoE: Communication-Aware Adaptive Routing with Load-Balanced Expert Parallelism for Efficient Mixture-of-Experts Training
von: Jin, Haopeng
Veröffentlicht: (2026)
von: Jin, Haopeng
Veröffentlicht: (2026)
Training Matryoshka Mixture-of-Experts for Elastic Inference-Time Expert Utilization
von: Wang, Yaoxiang, et al.
Veröffentlicht: (2025)
von: Wang, Yaoxiang, et al.
Veröffentlicht: (2025)
XL3M: A Training-free Framework for LLM Length Extension Based on Segment-wise Inference
von: Wang, Shengnan, et al.
Veröffentlicht: (2024)
von: Wang, Shengnan, et al.
Veröffentlicht: (2024)
Toward Efficient Inference Attacks: Shadow Model Sharing via Mixture-of-Experts
von: Bai, Li, et al.
Veröffentlicht: (2025)
von: Bai, Li, et al.
Veröffentlicht: (2025)
Efficient Long-Context LLM Inference via KV Cache Clustering
von: Hu, Jie, et al.
Veröffentlicht: (2025)
von: Hu, Jie, et al.
Veröffentlicht: (2025)
Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs
von: Liu, Enshu, et al.
Veröffentlicht: (2024)
von: Liu, Enshu, et al.
Veröffentlicht: (2024)
Efficient Mixture-of-Experts LLM Inference with Apple Silicon NPUs
von: Benazir, Afsara, et al.
Veröffentlicht: (2026)
von: Benazir, Afsara, et al.
Veröffentlicht: (2026)
Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models
von: Pan, Bowen, et al.
Veröffentlicht: (2024)
von: Pan, Bowen, et al.
Veröffentlicht: (2024)
HeterMoE: Efficient Training of Mixture-of-Experts Models on Heterogeneous GPUs
von: Wu, Yongji, et al.
Veröffentlicht: (2025)
von: Wu, Yongji, et al.
Veröffentlicht: (2025)
MoTE: Mixture of Task-specific Experts for Pre-Trained ModelBased Class-incremental Learning
von: Li, Linjie, et al.
Veröffentlicht: (2025)
von: Li, Linjie, et al.
Veröffentlicht: (2025)
Fate: Fast Edge Inference of Mixture-of-Experts Models via Cross-Layer Gate
von: Fang, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Fang, Zhiyuan, et al.
Veröffentlicht: (2025)
Uncovering Intra-expert Activation Sparsity for Efficient Mixture-of-Expert Model Execution
von: Park, Jongseok, et al.
Veröffentlicht: (2026)
von: Park, Jongseok, et al.
Veröffentlicht: (2026)
Lagom: Unleashing the Power of Communication and Computation Overlapping for Distributed LLM Training
von: Xu, Guanbin, et al.
Veröffentlicht: (2026)
von: Xu, Guanbin, et al.
Veröffentlicht: (2026)
Dynamic Mixture of Experts: An Auto-Tuning Approach for Efficient Transformer Models
von: Guo, Yongxin, et al.
Veröffentlicht: (2024)
von: Guo, Yongxin, et al.
Veröffentlicht: (2024)
Fast Training of Mixture-of-Experts for Time Series Forecasting via Expert Loss Integration
von: Mahtout, Btissame El, et al.
Veröffentlicht: (2026)
von: Mahtout, Btissame El, et al.
Veröffentlicht: (2026)
MoBiE: Efficient Inference of Mixture of Binary Experts under Post-Training Quantization
von: Zhao, Zhixiong, et al.
Veröffentlicht: (2026)
von: Zhao, Zhixiong, et al.
Veröffentlicht: (2026)
MoESys: A Distributed and Efficient Mixture-of-Experts Training and Inference System for Internet Services
von: Yu, Dianhai, et al.
Veröffentlicht: (2022)
von: Yu, Dianhai, et al.
Veröffentlicht: (2022)
Optimizing Mixture-of-Experts Inference Time Combining Model Deployment and Communication Scheduling
von: Li, Jialong, et al.
Veröffentlicht: (2024)
von: Li, Jialong, et al.
Veröffentlicht: (2024)
Elastic MoE: Unlocking the Inference-Time Scalability of Mixture-of-Experts
von: Gu, Naibin, et al.
Veröffentlicht: (2025)
von: Gu, Naibin, et al.
Veröffentlicht: (2025)
Lazarus: Resilient and Elastic Training of Mixture-of-Experts Models
von: Wu, Yongji, et al.
Veröffentlicht: (2024)
von: Wu, Yongji, et al.
Veröffentlicht: (2024)
Scalable Training of Mixture-of-Experts Models with Megatron Core
von: Yan, Zijie, et al.
Veröffentlicht: (2026)
von: Yan, Zijie, et al.
Veröffentlicht: (2026)
Efficient Training of Diffusion Mixture-of-Experts Models: A Practical Recipe
von: Liu, Yahui, et al.
Veröffentlicht: (2025)
von: Liu, Yahui, et al.
Veröffentlicht: (2025)
ExpertFlow: Efficient Mixture-of-Experts Inference via Predictive Expert Caching and Token Scheduling
von: He, Xin, et al.
Veröffentlicht: (2024)
von: He, Xin, et al.
Veröffentlicht: (2024)
LPT++: Efficient Training on Mixture of Long-tailed Experts
von: Dong, Bowen, et al.
Veröffentlicht: (2024)
von: Dong, Bowen, et al.
Veröffentlicht: (2024)
MELINOE: Fine-Tuning Enables Memory-Efficient Inference for Mixture-of-Experts Models
von: Raje, Arian, et al.
Veröffentlicht: (2026)
von: Raje, Arian, et al.
Veröffentlicht: (2026)
RailS: Load Balancing for All-to-All Communication in Distributed Mixture-of-Experts Training
von: Xu, Heng, et al.
Veröffentlicht: (2025)
von: Xu, Heng, et al.
Veröffentlicht: (2025)
Speculating Experts Accelerates Inference for Mixture-of-Experts
von: Madan, Vivan, et al.
Veröffentlicht: (2026)
von: Madan, Vivan, et al.
Veröffentlicht: (2026)
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding
von: Bang, Jehyeon, et al.
Veröffentlicht: (2026)
von: Bang, Jehyeon, et al.
Veröffentlicht: (2026)
Mixture of Cache-Conditional Experts for Efficient Mobile Device Inference
von: Skliar, Andrii, et al.
Veröffentlicht: (2024)
von: Skliar, Andrii, et al.
Veröffentlicht: (2024)
PreMoE: Proactive Inference for Efficient Mixture-of-Experts
von: Pei, Zehua, et al.
Veröffentlicht: (2025)
von: Pei, Zehua, et al.
Veröffentlicht: (2025)
Mixture of Style Experts for Diverse Image Stylization
von: Zhu, Shihao, et al.
Veröffentlicht: (2026)
von: Zhu, Shihao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
BigMac: Breaking the Pareto Frontier of Compute and Memory in Multimodal LLM Training
von: Zhang, Zili, et al.
Veröffentlicht: (2026) -
Accelerating Long-Tail Generation in Synchronous RLHF Training via Adaptive Tensor Parallelism
von: Zhao, Long, et al.
Veröffentlicht: (2026) -
LiteCache: A Query Similarity-Driven, GPU-Centric KVCache Subsystem for Efficient LLM Inference
von: Yi, Jiawei, et al.
Veröffentlicht: (2025) -
AdaCluster: Adaptive Query-Key Clustering for Sparse Attention in Video Generation
von: Tan, Haoyue, et al.
Veröffentlicht: (2026) -
MoBiLE: Efficient Mixture-of-Experts Inference on Consumer GPU with Mixture of Big Little Experts
von: Zhao, Yushu, et al.
Veröffentlicht: (2025)