SMoE: An Algorithm-System Co-Design for Pushing MoE to the Edge via Expert Substitution
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866909010825838592 |
|---|---|
| author | Zhu, Guoying Li, Meng Dai, Haipeng Liu, Xuechen Wang, Weijun Li, Keran xiao, Jun Chen, Ligeng Wang, Wei |
| author_facet | Zhu, Guoying Li, Meng Dai, Haipeng Liu, Xuechen Wang, Weijun Li, Keran xiao, Jun Chen, Ligeng Wang, Wei |
| contents | The Mixture of Experts (MoE) architecture has emerged as a key technique for scaling Large Language Models by activating only a subset of experts per query. Deploying MoE on consumer-grade edge hardware, however, is constrained by limited device memory, making dynamic expert offloading essential. Unlike prior work that treats offloading purely as a scheduling problem, we leverage expert importance to guide decisions, substituting low-importance activated experts with functionally similar ones already cached in GPU memory, thereby preserving accuracy. As a result, this design reduces memory usage and data transfer, while largely eliminating PCIe overhead. In addition, we introduce a scheduling policy that maximizes the reuse ratio of GPU-cached experts, further boosting efficiency. Extensive evaluations show that our approach delivers 48% lower decoding latency with over 60% expert cache hit rate, while maintaining nearly lossless accuracy. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_18983 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | SMoE: An Algorithm-System Co-Design for Pushing MoE to the Edge via Expert Substitution Zhu, Guoying Li, Meng Dai, Haipeng Liu, Xuechen Wang, Weijun Li, Keran xiao, Jun Chen, Ligeng Wang, Wei Artificial Intelligence The Mixture of Experts (MoE) architecture has emerged as a key technique for scaling Large Language Models by activating only a subset of experts per query. Deploying MoE on consumer-grade edge hardware, however, is constrained by limited device memory, making dynamic expert offloading essential. Unlike prior work that treats offloading purely as a scheduling problem, we leverage expert importance to guide decisions, substituting low-importance activated experts with functionally similar ones already cached in GPU memory, thereby preserving accuracy. As a result, this design reduces memory usage and data transfer, while largely eliminating PCIe overhead. In addition, we introduce a scheduling policy that maximizes the reuse ratio of GPU-cached experts, further boosting efficiency. Extensive evaluations show that our approach delivers 48% lower decoding latency with over 60% expert cache hit rate, while maintaining nearly lossless accuracy. |
| title | SMoE: An Algorithm-System Co-Design for Pushing MoE to the Edge via Expert Substitution |
| topic | Artificial Intelligence |
| url | https://arxiv.org/abs/2508.18983 |