Remoe: Towards Efficient and Low-Cost MoE Inference in Serverless Computing
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Liu, Wentao, Hu, Yuhao, Zhou, Ruiting, Li, Baochun, Wang, Ne |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
MoEless: Efficient MoE LLM Serving via Serverless Computing
par: Yu, Hanfei, et autres
Publié: (2026)
par: Yu, Hanfei, et autres
Publié: (2026)
D$^{2}$MoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM Serving
par: Wang, Haodong, et autres
Publié: (2025)
par: Wang, Haodong, et autres
Publié: (2025)
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems
par: Wu, Qi, et autres
Publié: (2026)
par: Wu, Qi, et autres
Publié: (2026)
Efficient MoE Inference with Fine-Grained Scheduling of Disaggregated Expert Parallelism
par: Pan, Xinglin, et autres
Publié: (2025)
par: Pan, Xinglin, et autres
Publié: (2025)
MoE-Lens: Towards the Hardware Limit of High-Throughput MoE LLM Serving Under Resource Constraints
par: Yuan, Yichao, et autres
Publié: (2025)
par: Yuan, Yichao, et autres
Publié: (2025)
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs
par: Cao, Shiyi, et autres
Publié: (2024)
par: Cao, Shiyi, et autres
Publié: (2024)
RLHFless: Serverless Computing for Efficient RLHF
par: Wei, Rui, et autres
Publié: (2026)
par: Wei, Rui, et autres
Publié: (2026)
ExpertFlow: Adaptive Expert Scheduling and Memory Coordination for Efficient MoE Inference
par: Shen, Zixu, et autres
Publié: (2025)
par: Shen, Zixu, et autres
Publié: (2025)
EPS-MoE: Expert Pipeline Scheduler for Cost-Efficient MoE Inference
par: Qian, Yulei, et autres
Publié: (2024)
par: Qian, Yulei, et autres
Publié: (2024)
Accelerating MoE Model Inference with Expert Sharding
par: Balmau, Oana, et autres
Publié: (2025)
par: Balmau, Oana, et autres
Publié: (2025)
Semantic Parallelism: Redefining Efficient MoE Inference via Model-Data Co-Scheduling
par: Li, Yan, et autres
Publié: (2025)
par: Li, Yan, et autres
Publié: (2025)
ProMoE: Fast MoE-based LLM Serving using Proactive Caching
par: Song, Xiaoniu, et autres
Publié: (2024)
par: Song, Xiaoniu, et autres
Publié: (2024)
MoEBlaze: Breaking the Memory Wall for Efficient MoE Training on Modern GPUs
par: Zhang, Jiyuan, et autres
Publié: (2026)
par: Zhang, Jiyuan, et autres
Publié: (2026)
ZipMoE: Efficient On-Device MoE Serving via Lossless Compression and Cache-Affinity Scheduling
par: Yang, Yuchen, et autres
Publié: (2026)
par: Yang, Yuchen, et autres
Publié: (2026)
ENOVA: Autoscaling towards Cost-effective and Stable Serverless LLM Serving
par: Huang, Tao, et autres
Publié: (2024)
par: Huang, Tao, et autres
Publié: (2024)
GRACE-MoE: Grouping and Replication with Locality-Aware Routing for Efficient Distributed MoE Inference
par: Han, Yu, et autres
Publié: (2025)
par: Han, Yu, et autres
Publié: (2025)
TriMoE: Augmenting GPU with AMX-Enabled CPU and DIMM-NDP for High-Throughput MoE Inference via Offloading
par: Pan, Yudong, et autres
Publié: (2026)
par: Pan, Yudong, et autres
Publié: (2026)
DisagMoE: Computation-Communication overlapped MoE Training via Disaggregated AF-Pipe Parallelism
par: Zeng, Zhichen, et autres
Publié: (2026)
par: Zeng, Zhichen, et autres
Publié: (2026)
Patterns behind Chaos: Forecasting Data Movement for Efficient Large-Scale MoE LLM Inference
par: Yu, Zhongkai, et autres
Publié: (2025)
par: Yu, Zhongkai, et autres
Publié: (2025)
Token-Budget-Aware Pool Routing for Cost-Efficient LLM Inference
par: Chen, Huamin, et autres
Publié: (2026)
par: Chen, Huamin, et autres
Publié: (2026)
OD-MoE: On-Demand Expert Loading for Cacheless Edge-Distributed MoE Inference
par: Wang, Liujianfu, et autres
Publié: (2025)
par: Wang, Liujianfu, et autres
Publié: (2025)
Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient Inference
par: Yu, Minchen, et autres
Publié: (2023)
par: Yu, Minchen, et autres
Publié: (2023)
Calibre: Towards Fair and Accurate Personalized Federated Learning with Self-Supervised Learning
par: Chen, Sijia, et autres
Publié: (2024)
par: Chen, Sijia, et autres
Publié: (2024)
HarMoEny: Efficient Multi-GPU Inference of MoE Models
par: Doucet, Zachary, et autres
Publié: (2025)
par: Doucet, Zachary, et autres
Publié: (2025)
SP-MoE: Speculative Decoding and Prefetching for Accelerating MoE-based Model Inference
par: Chen, Liangkun, et autres
Publié: (2025)
par: Chen, Liangkun, et autres
Publié: (2025)
Towards Resource-Efficient Serverless LLM Inference with SLINFER
par: Xu, Chuhao, et autres
Publié: (2025)
par: Xu, Chuhao, et autres
Publié: (2025)
AIBrix: Towards Scalable, Cost-Effective Large Language Model Inference Infrastructure
par: The AIBrix Team, et autres
Publié: (2025)
par: The AIBrix Team, et autres
Publié: (2025)
Accelerating Distributed MoE Training and Inference with Lina
par: Li, Jiamin, et autres
Publié: (2022)
par: Li, Jiamin, et autres
Publié: (2022)
Word Frequency Counting Based on Serverless MapReduce
par: Li, Hanzhe, et autres
Publié: (2026)
par: Li, Hanzhe, et autres
Publié: (2026)
Cluster Topology-Driven Placement of Experts Reduces Network Traffic in MoE Inference
par: Sivtsov, Danil, et autres
Publié: (2025)
par: Sivtsov, Danil, et autres
Publié: (2025)
Piper: Efficient Large-Scale MoE Training via Resource Modeling and Pipelined Hybrid Parallelism
par: Dash, Sajal, et autres
Publié: (2026)
par: Dash, Sajal, et autres
Publié: (2026)
FSD-Inference: Fully Serverless Distributed Inference with Scalable Cloud Communication
par: Oakley, Joe, et autres
Publié: (2024)
par: Oakley, Joe, et autres
Publié: (2024)
Efficient CPU-GPU Collaborative Inference for MoE-based LLMs on Memory-Limited Systems
par: Huang, En-Ming, et autres
Publié: (2025)
par: Huang, En-Ming, et autres
Publié: (2025)
MoESys: A Distributed and Efficient Mixture-of-Experts Training and Inference System for Internet Services
par: Yu, Dianhai, et autres
Publié: (2022)
par: Yu, Dianhai, et autres
Publié: (2022)
LSH-MoE: Communication-efficient MoE Training via Locality-Sensitive Hashing
par: Nie, Xiaonan, et autres
Publié: (2024)
par: Nie, Xiaonan, et autres
Publié: (2024)
PROBE: Co-Balancing Computation and Communication in MoE Inference via Real-Time Predictive Prefetching
par: Zhu, Qianchao, et autres
Publié: (2026)
par: Zhu, Qianchao, et autres
Publié: (2026)
Expert-as-a-Service: Towards Efficient, Scalable, and Robust Large-scale MoE Serving
par: Liu, Ziming, et autres
Publié: (2025)
par: Liu, Ziming, et autres
Publié: (2025)
Towards Energy-Efficient Serverless Computing with Hardware Isolation
par: Carl, Natalie, et autres
Publié: (2025)
par: Carl, Natalie, et autres
Publié: (2025)
Relay Buffer Independent Communication over Pooled HBM for Efficient MoE Inference on Ascend
par: Hu, Tianlun, et autres
Publié: (2026)
par: Hu, Tianlun, et autres
Publié: (2026)
Green or Fast? Learning to Balance Cold Starts and Idle Carbon in Serverless Computing
par: Sun, Bowen, et autres
Publié: (2026)
par: Sun, Bowen, et autres
Publié: (2026)
Documents similaires
-
MoEless: Efficient MoE LLM Serving via Serverless Computing
par: Yu, Hanfei, et autres
Publié: (2026) -
D$^{2}$MoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM Serving
par: Wang, Haodong, et autres
Publié: (2025) -
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems
par: Wu, Qi, et autres
Publié: (2026) -
Efficient MoE Inference with Fine-Grained Scheduling of Disaggregated Expert Parallelism
par: Pan, Xinglin, et autres
Publié: (2025) -
MoE-Lens: Towards the Hardware Limit of High-Throughput MoE LLM Serving Under Resource Constraints
par: Yuan, Yichao, et autres
Publié: (2025)