GEM: GPU-Variability-Aware Expert to GPU Mapping for MoE Systems
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wawdhane, Sourish, Kumar, Avinash, Das, Poulami |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems
von: Wu, Qi, et al.
Veröffentlicht: (2026)
von: Wu, Qi, et al.
Veröffentlicht: (2026)
Linear-MoE: Linear Sequence Modeling Meets Mixture-of-Experts
von: Sun, Weigao, et al.
Veröffentlicht: (2025)
von: Sun, Weigao, et al.
Veröffentlicht: (2025)
EPS-MoE: Expert Pipeline Scheduler for Cost-Efficient MoE Inference
von: Qian, Yulei, et al.
Veröffentlicht: (2024)
von: Qian, Yulei, et al.
Veröffentlicht: (2024)
LeMix: Unified Scheduling for LLM Training and Inference on Multi-GPU Systems
von: Li, Yufei, et al.
Veröffentlicht: (2025)
von: Li, Yufei, et al.
Veröffentlicht: (2025)
Efficient MoE Inference with Fine-Grained Scheduling of Disaggregated Expert Parallelism
von: Pan, Xinglin, et al.
Veröffentlicht: (2025)
von: Pan, Xinglin, et al.
Veröffentlicht: (2025)
HarMoEny: Efficient Multi-GPU Inference of MoE Models
von: Doucet, Zachary, et al.
Veröffentlicht: (2025)
von: Doucet, Zachary, et al.
Veröffentlicht: (2025)
TriMoE: Augmenting GPU with AMX-Enabled CPU and DIMM-NDP for High-Throughput MoE Inference via Offloading
von: Pan, Yudong, et al.
Veröffentlicht: (2026)
von: Pan, Yudong, et al.
Veröffentlicht: (2026)
GPU-Virt-Bench: A Comprehensive Benchmarking Framework for Software-Based GPU Virtualization Systems
von: VG, Jithin, et al.
Veröffentlicht: (2025)
von: VG, Jithin, et al.
Veröffentlicht: (2025)
ExpertFlow: Adaptive Expert Scheduling and Memory Coordination for Efficient MoE Inference
von: Shen, Zixu, et al.
Veröffentlicht: (2025)
von: Shen, Zixu, et al.
Veröffentlicht: (2025)
Efficient CPU-GPU Collaborative Inference for MoE-based LLMs on Memory-Limited Systems
von: Huang, En-Ming, et al.
Veröffentlicht: (2025)
von: Huang, En-Ming, et al.
Veröffentlicht: (2025)
PithTrain: A Compact and Agent-Native MoE Training System
von: Lai, Ruihang, et al.
Veröffentlicht: (2026)
von: Lai, Ruihang, et al.
Veröffentlicht: (2026)
D$^{2}$MoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM Serving
von: Wang, Haodong, et al.
Veröffentlicht: (2025)
von: Wang, Haodong, et al.
Veröffentlicht: (2025)
MoE-Lens: Towards the Hardware Limit of High-Throughput MoE LLM Serving Under Resource Constraints
von: Yuan, Yichao, et al.
Veröffentlicht: (2025)
von: Yuan, Yichao, et al.
Veröffentlicht: (2025)
Accelerating MoE Model Inference with Expert Sharding
von: Balmau, Oana, et al.
Veröffentlicht: (2025)
von: Balmau, Oana, et al.
Veröffentlicht: (2025)
Advancing MoE Efficiency: A Collaboration-Constrained Routing (C2R) Strategy for Better Expert Parallelism Design
von: Zhang, Mohan, et al.
Veröffentlicht: (2025)
von: Zhang, Mohan, et al.
Veröffentlicht: (2025)
ARGUS: Agentic GPU Optimization Guided by Data-Flow Invariants
von: Mai, Haohui, et al.
Veröffentlicht: (2026)
von: Mai, Haohui, et al.
Veröffentlicht: (2026)
SwizzlePerf: Hardware-Aware LLMs for GPU Kernel Performance Optimization
von: Tschand, Arya, et al.
Veröffentlicht: (2025)
von: Tschand, Arya, et al.
Veröffentlicht: (2025)
ProMoE: Fast MoE-based LLM Serving using Proactive Caching
von: Song, Xiaoniu, et al.
Veröffentlicht: (2024)
von: Song, Xiaoniu, et al.
Veröffentlicht: (2024)
Breaking the MoE LLM Trilemma: Dynamic Expert Clustering with Structured Compression
von: Zhu, Peijun, et al.
Veröffentlicht: (2025)
von: Zhu, Peijun, et al.
Veröffentlicht: (2025)
Watt Counts: Energy-Aware Benchmark for Sustainable LLM Inference on Heterogeneous GPU Architectures
von: Argerich, Mauricio Fadel, et al.
Veröffentlicht: (2026)
von: Argerich, Mauricio Fadel, et al.
Veröffentlicht: (2026)
MoE-Hub: Taming Software Complexity for Seamless MoE Overlap with Hardware-Accelerated Communication on Multi-GPU Systems
von: Zhou, Zhuoshan, et al.
Veröffentlicht: (2026)
von: Zhou, Zhuoshan, et al.
Veröffentlicht: (2026)
GPUTOK: GPU Accelerated Byte Level BPE Tokenization
von: Kadamba, Venu Gopal, et al.
Veröffentlicht: (2026)
von: Kadamba, Venu Gopal, et al.
Veröffentlicht: (2026)
MoE-Gen: High-Throughput MoE Inference on a Single GPU with Module-Based Batching
von: Xu, Tairan, et al.
Veröffentlicht: (2025)
von: Xu, Tairan, et al.
Veröffentlicht: (2025)
Remoe: Towards Efficient and Low-Cost MoE Inference in Serverless Computing
von: Liu, Wentao, et al.
Veröffentlicht: (2025)
von: Liu, Wentao, et al.
Veröffentlicht: (2025)
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs
von: Cao, Shiyi, et al.
Veröffentlicht: (2024)
von: Cao, Shiyi, et al.
Veröffentlicht: (2024)
UCCL-Zip: Lossless Compression Supercharged GPU Communication
von: Ma, Shuang, et al.
Veröffentlicht: (2026)
von: Ma, Shuang, et al.
Veröffentlicht: (2026)
MSCCL++: Rethinking GPU Communication Abstractions for AI Inference
von: Hwang, Changho, et al.
Veröffentlicht: (2025)
von: Hwang, Changho, et al.
Veröffentlicht: (2025)
Power- and Fragmentation-aware Online Scheduling for GPU Datacenters
von: Lettich, Francesco, et al.
Veröffentlicht: (2024)
von: Lettich, Francesco, et al.
Veröffentlicht: (2024)
OD-MoE: On-Demand Expert Loading for Cacheless Edge-Distributed MoE Inference
von: Wang, Liujianfu, et al.
Veröffentlicht: (2025)
von: Wang, Liujianfu, et al.
Veröffentlicht: (2025)
Data Driven Optimization of GPU efficiency for Distributed LLM Adapter Serving
von: Agullo, Ferran, et al.
Veröffentlicht: (2026)
von: Agullo, Ferran, et al.
Veröffentlicht: (2026)
Beyond the GPU: The Strategic Role of FPGAs in the Next Wave of AI
von: Jiménez, Arturo Urías
Veröffentlicht: (2025)
von: Jiménez, Arturo Urías
Veröffentlicht: (2025)
An Efficient Heterogeneous Co-Design for Fine-Tuning on a Single GPU
von: Yang, Ruijia, et al.
Veröffentlicht: (2026)
von: Yang, Ruijia, et al.
Veröffentlicht: (2026)
Towards Scalable GPU-Accelerated SNN Training via Temporal Fusion
von: Li, Yanchen, et al.
Veröffentlicht: (2024)
von: Li, Yanchen, et al.
Veröffentlicht: (2024)
Accelerating Large Language Model Training with Hybrid GPU-based Compression
von: Xu, Lang, et al.
Veröffentlicht: (2024)
von: Xu, Lang, et al.
Veröffentlicht: (2024)
HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE Inference
von: Zhong, Shuzhang, et al.
Veröffentlicht: (2025)
von: Zhong, Shuzhang, et al.
Veröffentlicht: (2025)
FourierMoE: Fourier Mixture-of-Experts Adaptation of Large Language Models
von: Jiang, Juyong, et al.
Veröffentlicht: (2026)
von: Jiang, Juyong, et al.
Veröffentlicht: (2026)
OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models
von: Xue, Fuzhao, et al.
Veröffentlicht: (2024)
von: Xue, Fuzhao, et al.
Veröffentlicht: (2024)
Xe-Forge: Multi-Stage LLM-Powered Kernel Optimization for Intel GPU
von: Spoczynski, Marcin, et al.
Veröffentlicht: (2026)
von: Spoczynski, Marcin, et al.
Veröffentlicht: (2026)
Lobster: A GPU-Accelerated Framework for Neurosymbolic Programming
von: Biberstein, Paul, et al.
Veröffentlicht: (2025)
von: Biberstein, Paul, et al.
Veröffentlicht: (2025)
Speeding up Local Optimization in Vehicle Routing with Tensor-based GPU Acceleration
von: Lei, Zhenyu, et al.
Veröffentlicht: (2025)
von: Lei, Zhenyu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems
von: Wu, Qi, et al.
Veröffentlicht: (2026) -
Linear-MoE: Linear Sequence Modeling Meets Mixture-of-Experts
von: Sun, Weigao, et al.
Veröffentlicht: (2025) -
EPS-MoE: Expert Pipeline Scheduler for Cost-Efficient MoE Inference
von: Qian, Yulei, et al.
Veröffentlicht: (2024) -
LeMix: Unified Scheduling for LLM Training and Inference on Multi-GPU Systems
von: Li, Yufei, et al.
Veröffentlicht: (2025) -
Efficient MoE Inference with Fine-Grained Scheduling of Disaggregated Expert Parallelism
von: Pan, Xinglin, et al.
Veröffentlicht: (2025)