Accelerating Mixture-of-Experts Inference by Hiding Offloading Latency with Speculative Decoding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Zhibin, Zhang, Zhonghui, Zhou, Yuhang, Wang, Zibo, Zhou, Mo, Jiang, Peng, Cai, Weilin, Huan, Chengying, Gu, Rong, Zhong, Sheng, Tian, Chen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!