HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE Inference
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhong, Shuzhang, Sun, Yanfan, Liang, Ling, Wang, Runsheng, Huang, Ru, Li, Meng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Efficient CPU-GPU Collaborative Inference for MoE-based LLMs on Memory-Limited Systems
von: Huang, En-Ming, et al.
Veröffentlicht: (2025)
von: Huang, En-Ming, et al.
Veröffentlicht: (2025)
HarMoEny: Efficient Multi-GPU Inference of MoE Models
von: Doucet, Zachary, et al.
Veröffentlicht: (2025)
von: Doucet, Zachary, et al.
Veröffentlicht: (2025)
EPS-MoE: Expert Pipeline Scheduler for Cost-Efficient MoE Inference
von: Qian, Yulei, et al.
Veröffentlicht: (2024)
von: Qian, Yulei, et al.
Veröffentlicht: (2024)
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems
von: Wu, Qi, et al.
Veröffentlicht: (2026)
von: Wu, Qi, et al.
Veröffentlicht: (2026)
GRACE-MoE: Grouping and Replication with Locality-Aware Routing for Efficient Distributed MoE Inference
von: Han, Yu, et al.
Veröffentlicht: (2025)
von: Han, Yu, et al.
Veröffentlicht: (2025)
SP-MoE: Speculative Decoding and Prefetching for Accelerating MoE-based Model Inference
von: Chen, Liangkun, et al.
Veröffentlicht: (2025)
von: Chen, Liangkun, et al.
Veröffentlicht: (2025)
OD-MoE: On-Demand Expert Loading for Cacheless Edge-Distributed MoE Inference
von: Wang, Liujianfu, et al.
Veröffentlicht: (2025)
von: Wang, Liujianfu, et al.
Veröffentlicht: (2025)
MoE-Gen: High-Throughput MoE Inference on a Single GPU with Module-Based Batching
von: Xu, Tairan, et al.
Veröffentlicht: (2025)
von: Xu, Tairan, et al.
Veröffentlicht: (2025)
Multi-Layer Scheduling for MoE-Based LLM Reasoning
von: Sun, Yifan, et al.
Veröffentlicht: (2026)
von: Sun, Yifan, et al.
Veröffentlicht: (2026)
ProMoE: Fast MoE-based LLM Serving using Proactive Caching
von: Song, Xiaoniu, et al.
Veröffentlicht: (2024)
von: Song, Xiaoniu, et al.
Veröffentlicht: (2024)
TriMoE: Augmenting GPU with AMX-Enabled CPU and DIMM-NDP for High-Throughput MoE Inference via Offloading
von: Pan, Yudong, et al.
Veröffentlicht: (2026)
von: Pan, Yudong, et al.
Veröffentlicht: (2026)
Efficient MoE Inference with Fine-Grained Scheduling of Disaggregated Expert Parallelism
von: Pan, Xinglin, et al.
Veröffentlicht: (2025)
von: Pan, Xinglin, et al.
Veröffentlicht: (2025)
MemFine: Memory-Aware Fine-Grained Scheduling for MoE Training
von: Zhao, Lu, et al.
Veröffentlicht: (2025)
von: Zhao, Lu, et al.
Veröffentlicht: (2025)
ZipMoE: Efficient On-Device MoE Serving via Lossless Compression and Cache-Affinity Scheduling
von: Yang, Yuchen, et al.
Veröffentlicht: (2026)
von: Yang, Yuchen, et al.
Veröffentlicht: (2026)
DuoServe-MoE: Dual-Phase Expert Prefetch and Caching for LLM Inference QoS Assurance
von: Zhang, Yuning, et al.
Veröffentlicht: (2025)
von: Zhang, Yuning, et al.
Veröffentlicht: (2025)
Accelerating Distributed MoE Training and Inference with Lina
von: Li, Jiamin, et al.
Veröffentlicht: (2022)
von: Li, Jiamin, et al.
Veröffentlicht: (2022)
D$^{2}$MoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM Serving
von: Wang, Haodong, et al.
Veröffentlicht: (2025)
von: Wang, Haodong, et al.
Veröffentlicht: (2025)
MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism
von: Zhang, Zheng, et al.
Veröffentlicht: (2025)
von: Zhang, Zheng, et al.
Veröffentlicht: (2025)
LSH-MoE: Communication-efficient MoE Training via Locality-Sensitive Hashing
von: Nie, Xiaonan, et al.
Veröffentlicht: (2024)
von: Nie, Xiaonan, et al.
Veröffentlicht: (2024)
Staleness-Centric Optimizations for Parallel Diffusion MoE Inference
von: Luo, Jiajun, et al.
Veröffentlicht: (2024)
von: Luo, Jiajun, et al.
Veröffentlicht: (2024)
Janus: Disaggregating Attention and Experts for Scalable MoE Inference
von: Zhang, Zhexiang, et al.
Veröffentlicht: (2025)
von: Zhang, Zhexiang, et al.
Veröffentlicht: (2025)
MoE-Hub: Taming Software Complexity for Seamless MoE Overlap with Hardware-Accelerated Communication on Multi-GPU Systems
von: Zhou, Zhuoshan, et al.
Veröffentlicht: (2026)
von: Zhou, Zhuoshan, et al.
Veröffentlicht: (2026)
Priority-Aware Preemptive Scheduling for Mixed-Priority Workloads in MoE Inference
von: Siavashi, Mohammad, et al.
Veröffentlicht: (2025)
von: Siavashi, Mohammad, et al.
Veröffentlicht: (2025)
ReviveMoE: Fast Recovery for Hardware Failures in Large-Scale MoE LLM Inference Deployments
von: Li, Haley, et al.
Veröffentlicht: (2026)
von: Li, Haley, et al.
Veröffentlicht: (2026)
ExpertFlow: Adaptive Expert Scheduling and Memory Coordination for Efficient MoE Inference
von: Shen, Zixu, et al.
Veröffentlicht: (2025)
von: Shen, Zixu, et al.
Veröffentlicht: (2025)
MoE-Compression: How the Compression Error of Experts Affects the Inference Accuracy of MoE Model?
von: Ma, Songkai, et al.
Veröffentlicht: (2025)
von: Ma, Songkai, et al.
Veröffentlicht: (2025)
Surviving Partial Rank Failures in Wide Expert-Parallel MoE Inference
von: Sun, Xun, et al.
Veröffentlicht: (2026)
von: Sun, Xun, et al.
Veröffentlicht: (2026)
ReaLB: Real-Time Load Balancing for Multimodal MoE Inference
von: Wang, Yingping, et al.
Veröffentlicht: (2026)
von: Wang, Yingping, et al.
Veröffentlicht: (2026)
eMoE: Task-aware Memory Efficient Mixture-of-Experts-Based (MoE) Model Inference
von: Tairin, Suraiya, et al.
Veröffentlicht: (2025)
von: Tairin, Suraiya, et al.
Veröffentlicht: (2025)
Hexa-MoE: Efficient and Heterogeneous-aware Training for Mixture-of-Experts
von: Luo, Shuqing, et al.
Veröffentlicht: (2024)
von: Luo, Shuqing, et al.
Veröffentlicht: (2024)
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs
von: Cao, Shiyi, et al.
Veröffentlicht: (2024)
von: Cao, Shiyi, et al.
Veröffentlicht: (2024)
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
von: Lin, Mao, et al.
Veröffentlicht: (2026)
von: Lin, Mao, et al.
Veröffentlicht: (2026)
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement
von: Wu, Tian, et al.
Veröffentlicht: (2025)
von: Wu, Tian, et al.
Veröffentlicht: (2025)
Remoe: Towards Efficient and Low-Cost MoE Inference in Serverless Computing
von: Liu, Wentao, et al.
Veröffentlicht: (2025)
von: Liu, Wentao, et al.
Veröffentlicht: (2025)
EC2MoE: Adaptive End-Cloud Pipeline Collaboration Enabling Scalable Mixture-of-Experts Inference
von: Yang, Zheming, et al.
Veröffentlicht: (2025)
von: Yang, Zheming, et al.
Veröffentlicht: (2025)
Sparse Checkpointing for Fast and Reliable MoE Training
von: Gandhi, Swapnil, et al.
Veröffentlicht: (2024)
von: Gandhi, Swapnil, et al.
Veröffentlicht: (2024)
DALI: A Workload-Aware Offloading Framework for Efficient MoE Inference on Local PCs
von: Zhu, Zeyu, et al.
Veröffentlicht: (2026)
von: Zhu, Zeyu, et al.
Veröffentlicht: (2026)
Semantic Parallelism: Redefining Efficient MoE Inference via Model-Data Co-Scheduling
von: Li, Yan, et al.
Veröffentlicht: (2025)
von: Li, Yan, et al.
Veröffentlicht: (2025)
Making MoE-based LLM Inference Resilient with Tarragon
von: Zhang, Songyu, et al.
Veröffentlicht: (2026)
von: Zhang, Songyu, et al.
Veröffentlicht: (2026)
Relay Buffer Independent Communication over Pooled HBM for Efficient MoE Inference on Ascend
von: Hu, Tianlun, et al.
Veröffentlicht: (2026)
von: Hu, Tianlun, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Efficient CPU-GPU Collaborative Inference for MoE-based LLMs on Memory-Limited Systems
von: Huang, En-Ming, et al.
Veröffentlicht: (2025) -
HarMoEny: Efficient Multi-GPU Inference of MoE Models
von: Doucet, Zachary, et al.
Veröffentlicht: (2025) -
EPS-MoE: Expert Pipeline Scheduler for Cost-Efficient MoE Inference
von: Qian, Yulei, et al.
Veröffentlicht: (2024) -
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems
von: Wu, Qi, et al.
Veröffentlicht: (2026) -
GRACE-MoE: Grouping and Replication with Locality-Aware Routing for Efficient Distributed MoE Inference
von: Han, Yu, et al.
Veröffentlicht: (2025)