WDMoE: Wireless Distributed Mixture of Experts for Large Language Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Xue, Nan, Sun, Yaping, Chen, Zhiyong, Tao, Meixia, Xu, Xiaodong, Qian, Liang, Cui, Shuguang, Zhang, Wenjun, Zhang, Ping |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
WDMoE: Wireless Distributed Large Language Models with Mixture of Experts
por: Xue, Nan, et al.
Publicado: (2024)
por: Xue, Nan, et al.
Publicado: (2024)
Optimal Expert Selection for Distributed Mixture-of-Experts at the Wireless Edge
por: Qin, Shengling, et al.
Publicado: (2025)
por: Qin, Shengling, et al.
Publicado: (2025)
Stable-MoE: Lyapunov-based Token Routing for Distributed Mixture-of-Experts Training over Edge Networks
por: Shi, Long, et al.
Publicado: (2025)
por: Shi, Long, et al.
Publicado: (2025)
FlowMoE: A Scalable Pipeline Scheduling Framework for Distributed Mixture-of-Experts Training
por: Gao, Yunqi, et al.
Publicado: (2025)
por: Gao, Yunqi, et al.
Publicado: (2025)
ElasticMoE: An Efficient Auto Scaling Method for Mixture-of-Experts Models
por: Singh, Gursimran, et al.
Publicado: (2025)
por: Singh, Gursimran, et al.
Publicado: (2025)
Split Fine-Tuning for Large Language Models in Wireless Networks
por: Zhang, Songge, et al.
Publicado: (2025)
por: Zhang, Songge, et al.
Publicado: (2025)
SplitLLM: Hierarchical Split Learning for Large Language Model over Wireless Network
por: Zhang, Songge, et al.
Publicado: (2025)
por: Zhang, Songge, et al.
Publicado: (2025)
Accelerating Mixture-of-Experts Inference by Hiding Offloading Latency with Speculative Decoding
por: Wang, Zhibin, et al.
Publicado: (2025)
por: Wang, Zhibin, et al.
Publicado: (2025)
ExpertWeave: Efficiently Serving Expert-Specialized Fine-Tuned Adapters at Scale
por: Shi, Ge, et al.
Publicado: (2025)
por: Shi, Ge, et al.
Publicado: (2025)
Toward Cost-Efficient Serving of Mixture-of-Experts with Asynchrony
por: Wang, Shaoyu, et al.
Publicado: (2025)
por: Wang, Shaoyu, et al.
Publicado: (2025)
HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference
por: Lin, Haoran, et al.
Publicado: (2025)
por: Lin, Haoran, et al.
Publicado: (2025)
Hexa-MoE: Efficient and Heterogeneous-aware Training for Mixture-of-Experts
por: Luo, Shuqing, et al.
Publicado: (2024)
por: Luo, Shuqing, et al.
Publicado: (2024)
Efficient Training of Large Language Models on Distributed Infrastructures: A Survey
por: Duan, Jiangfei, et al.
Publicado: (2024)
por: Duan, Jiangfei, et al.
Publicado: (2024)
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement
por: Wu, Tian, et al.
Publicado: (2025)
por: Wu, Tian, et al.
Publicado: (2025)
MoESys: A Distributed and Efficient Mixture-of-Experts Training and Inference System for Internet Services
por: Yu, Dianhai, et al.
Publicado: (2022)
por: Yu, Dianhai, et al.
Publicado: (2022)
Elastic Mixture of Rank-Wise Experts for Knowledge Reuse in Federated Fine-Tuning
por: Wu, Yebo, et al.
Publicado: (2025)
por: Wu, Yebo, et al.
Publicado: (2025)
DuoServe-MoE: Dual-Phase Expert Prefetch and Caching for LLM Inference QoS Assurance
por: Zhang, Yuning, et al.
Publicado: (2025)
por: Zhang, Yuning, et al.
Publicado: (2025)
Lazarus: Resilient and Elastic Training of Mixture-of-Experts Models
por: Wu, Yongji, et al.
Publicado: (2024)
por: Wu, Yongji, et al.
Publicado: (2024)
ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production
por: Xiang, Yuxing, et al.
Publicado: (2025)
por: Xiang, Yuxing, et al.
Publicado: (2025)
Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing
por: Liu, Mengfan, et al.
Publicado: (2025)
por: Liu, Mengfan, et al.
Publicado: (2025)
LAER-MoE: Load-Adaptive Expert Re-layout for Efficient Mixture-of-Experts Training
por: Liu, Xinyi, et al.
Publicado: (2026)
por: Liu, Xinyi, et al.
Publicado: (2026)
Low-Latency Federated Fine-Tuning for Large Language Models Over Wireless Networks
por: Pang, Zhiwen, et al.
Publicado: (2026)
por: Pang, Zhiwen, et al.
Publicado: (2026)
MoE-CAP: Benchmarking Cost, Accuracy and Performance of Sparse Mixture-of-Experts Systems
por: Jiang, Yinsicheng, et al.
Publicado: (2025)
por: Jiang, Yinsicheng, et al.
Publicado: (2025)
MoE-CAP: Benchmarking Cost, Accuracy and Performance of Sparse Mixture-of-Experts Systems
por: Jiang, Yinsicheng, et al.
Publicado: (2024)
por: Jiang, Yinsicheng, et al.
Publicado: (2024)
MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism
por: Zhu, Ruidong, et al.
Publicado: (2025)
por: Zhu, Ruidong, et al.
Publicado: (2025)
Scattered Mixture-of-Experts Implementation
por: Tan, Shawn, et al.
Publicado: (2024)
por: Tan, Shawn, et al.
Publicado: (2024)
Enhancing Convergence, Privacy and Fairness for Wireless Personalized Federated Learning: Quantization-Assisted Min-Max Fair Scheduling
por: Zhao, Xiyu, et al.
Publicado: (2025)
por: Zhao, Xiyu, et al.
Publicado: (2025)
Expert-as-a-Service: Towards Efficient, Scalable, and Robust Large-scale MoE Serving
por: Liu, Ziming, et al.
Publicado: (2025)
por: Liu, Ziming, et al.
Publicado: (2025)
SlimCaching: Edge Caching of Mixture-of-Experts for Distributed Inference
por: Chen, Qian, et al.
Publicado: (2025)
por: Chen, Qian, et al.
Publicado: (2025)
EC2MoE: Adaptive End-Cloud Pipeline Collaboration Enabling Scalable Mixture-of-Experts Inference
por: Yang, Zheming, et al.
Publicado: (2025)
por: Yang, Zheming, et al.
Publicado: (2025)
D-CAST: Distributed Consensus Switch in Wireless Trustworthy Autonomous System
por: Yu, Dachao, et al.
Publicado: (2024)
por: Yu, Dachao, et al.
Publicado: (2024)
A Novel Indicator for Quantifying and Minimizing Information Utility Loss of Robot Teams
por: Zhao, Xiyu, et al.
Publicado: (2025)
por: Zhao, Xiyu, et al.
Publicado: (2025)
Characterizing Communication Patterns in Distributed Large Language Model Inference
por: Xu, Lang, et al.
Publicado: (2025)
por: Xu, Lang, et al.
Publicado: (2025)
MoE-SpeQ: Speculative Quantized Decoding with Proactive Expert Prefetching and Offloading for Mixture-of-Experts
por: Wang, Wenfeng, et al.
Publicado: (2025)
por: Wang, Wenfeng, et al.
Publicado: (2025)
Towards the Distributed Large-scale k-NN Graph Construction by Graph Merge
por: Zhang, Cheng, et al.
Publicado: (2025)
por: Zhang, Cheng, et al.
Publicado: (2025)
Shortcut-connected Expert Parallelism for Accelerating Mixture-of-Experts
por: Cai, Weilin, et al.
Publicado: (2024)
por: Cai, Weilin, et al.
Publicado: (2024)
Optimal Transport Aggregation for Distributed Mixture-of-Experts
por: Chamroukhi, Faïcel, et al.
Publicado: (2023)
por: Chamroukhi, Faïcel, et al.
Publicado: (2023)
Gradient Compression and Correlation Driven Federated Learning for Wireless Traffic Prediction
por: Zhang, Chuanting, et al.
Publicado: (2025)
por: Zhang, Chuanting, et al.
Publicado: (2025)
Mitigating Interference of Microservices with a Scoring Mechanism in Large-scale Clusters
por: Yang, Dingyu, et al.
Publicado: (2024)
por: Yang, Dingyu, et al.
Publicado: (2024)
Minder: Faulty Machine Detection for Large-scale Distributed Model Training
por: Deng, Yangtao, et al.
Publicado: (2024)
por: Deng, Yangtao, et al.
Publicado: (2024)
Ejemplares similares
-
WDMoE: Wireless Distributed Large Language Models with Mixture of Experts
por: Xue, Nan, et al.
Publicado: (2024) -
Optimal Expert Selection for Distributed Mixture-of-Experts at the Wireless Edge
por: Qin, Shengling, et al.
Publicado: (2025) -
Stable-MoE: Lyapunov-based Token Routing for Distributed Mixture-of-Experts Training over Edge Networks
por: Shi, Long, et al.
Publicado: (2025) -
FlowMoE: A Scalable Pipeline Scheduling Framework for Distributed Mixture-of-Experts Training
por: Gao, Yunqi, et al.
Publicado: (2025) -
ElasticMoE: An Efficient Auto Scaling Method for Mixture-of-Experts Models
por: Singh, Gursimran, et al.
Publicado: (2025)