TriMoE: Augmenting GPU with AMX-Enabled CPU and DIMM-NDP for High-Throughput MoE Inference via Offloading

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Pan, Yudong, He, Yintao, Han, Tianhua, Liu, Lian, Zhao, Shixin, Chen, Zhirong, Wang, Mengdi, Li, Cangyuan, Han, Yinhe, Wang, Ying
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912934617153536
author Pan, Yudong
He, Yintao
Han, Tianhua
Liu, Lian
Zhao, Shixin
Chen, Zhirong
Wang, Mengdi
Li, Cangyuan
Han, Yinhe
Wang, Ying
author_facet Pan, Yudong
He, Yintao
Han, Tianhua
Liu, Lian
Zhao, Shixin
Chen, Zhirong
Wang, Mengdi
Li, Cangyuan
Han, Yinhe
Wang, Ying
contents To deploy large Mixture-of-Experts (MoE) models cost-effectively, offloading-based single-GPU heterogeneous inference is crucial. While GPU-CPU architectures that offload cold experts are constrained by host memory bandwidth, emerging GPU-NDP architectures utilize DIMM-NDP to offload non-hot experts. However, non-hot experts are not a homogeneous memory-bound group: a significant subset of warm experts exists is severely penalized by high GPU I/O latency yet can saturate NDP compute throughput, exposing a critical compute gap. We present TriMoE, a novel GPU-CPU-NDP architecture that fills this gap by synergistically leveraging AMX-enabled CPU to precisely map hot, warm, and cold experts onto their optimal compute units. We further introduce a bottleneck-aware expert scheduling policy and a prediction-driven dynamic relayout/rebalancing scheme. Experiments demonstrate that TriMoE achieves up to 2.83x speedup over state-of-the-art solutions.
format Preprint
id arxiv_https___arxiv_org_abs_2603_01058
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle TriMoE: Augmenting GPU with AMX-Enabled CPU and DIMM-NDP for High-Throughput MoE Inference via Offloading
Pan, Yudong
He, Yintao
Han, Tianhua
Liu, Lian
Zhao, Shixin
Chen, Zhirong
Wang, Mengdi
Li, Cangyuan
Han, Yinhe
Wang, Ying
Hardware Architecture
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
To deploy large Mixture-of-Experts (MoE) models cost-effectively, offloading-based single-GPU heterogeneous inference is crucial. While GPU-CPU architectures that offload cold experts are constrained by host memory bandwidth, emerging GPU-NDP architectures utilize DIMM-NDP to offload non-hot experts. However, non-hot experts are not a homogeneous memory-bound group: a significant subset of warm experts exists is severely penalized by high GPU I/O latency yet can saturate NDP compute throughput, exposing a critical compute gap. We present TriMoE, a novel GPU-CPU-NDP architecture that fills this gap by synergistically leveraging AMX-enabled CPU to precisely map hot, warm, and cold experts onto their optimal compute units. We further introduce a bottleneck-aware expert scheduling policy and a prediction-driven dynamic relayout/rebalancing scheme. Experiments demonstrate that TriMoE achieves up to 2.83x speedup over state-of-the-art solutions.
title TriMoE: Augmenting GPU with AMX-Enabled CPU and DIMM-NDP for High-Throughput MoE Inference via Offloading
topic Hardware Architecture
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2603.01058