MoE-Inference-Bench: Performance Evaluation of Mixture of Expert Large Language and Vision Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chitty-Venkata, Krishna Teja, Howland, Sylvia, Azar, Golara, Soboleva, Daria, Vassilieva, Natalia, Raskar, Siddhisanket, Emani, Murali, Vishwanath, Venkatram
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909751244226560
author Chitty-Venkata, Krishna Teja
Howland, Sylvia
Azar, Golara
Soboleva, Daria
Vassilieva, Natalia
Raskar, Siddhisanket
Emani, Murali
Vishwanath, Venkatram
author_facet Chitty-Venkata, Krishna Teja
Howland, Sylvia
Azar, Golara
Soboleva, Daria
Vassilieva, Natalia
Raskar, Siddhisanket
Emani, Murali
Vishwanath, Venkatram
contents Mixture of Experts (MoE) models have enabled the scaling of Large Language Models (LLMs) and Vision Language Models (VLMs) by achieving massive parameter counts while maintaining computational efficiency. However, MoEs introduce several inference-time challenges, including load imbalance across experts and the additional routing computational overhead. To address these challenges and fully harness the benefits of MoE, a systematic evaluation of hardware acceleration techniques is essential. We present MoE-Inference-Bench, a comprehensive study to evaluate MoE performance across diverse scenarios. We analyze the impact of batch size, sequence length, and critical MoE hyperparameters such as FFN dimensions and number of experts on throughput. We evaluate several optimization techniques on Nvidia H100 GPUs, including pruning, Fused MoE operations, speculative decoding, quantization, and various parallelization strategies. Our evaluation includes MoEs from the Mixtral, DeepSeek, OLMoE and Qwen families. The results reveal performance differences across configurations and provide insights for the efficient deployment of MoEs.
format Preprint
id arxiv_https___arxiv_org_abs_2508_17467
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MoE-Inference-Bench: Performance Evaluation of Mixture of Expert Large Language and Vision Models
Chitty-Venkata, Krishna Teja
Howland, Sylvia
Azar, Golara
Soboleva, Daria
Vassilieva, Natalia
Raskar, Siddhisanket
Emani, Murali
Vishwanath, Venkatram
Machine Learning
Performance
Mixture of Experts (MoE) models have enabled the scaling of Large Language Models (LLMs) and Vision Language Models (VLMs) by achieving massive parameter counts while maintaining computational efficiency. However, MoEs introduce several inference-time challenges, including load imbalance across experts and the additional routing computational overhead. To address these challenges and fully harness the benefits of MoE, a systematic evaluation of hardware acceleration techniques is essential. We present MoE-Inference-Bench, a comprehensive study to evaluate MoE performance across diverse scenarios. We analyze the impact of batch size, sequence length, and critical MoE hyperparameters such as FFN dimensions and number of experts on throughput. We evaluate several optimization techniques on Nvidia H100 GPUs, including pruning, Fused MoE operations, speculative decoding, quantization, and various parallelization strategies. Our evaluation includes MoEs from the Mixtral, DeepSeek, OLMoE and Qwen families. The results reveal performance differences across configurations and provide insights for the efficient deployment of MoEs.
title MoE-Inference-Bench: Performance Evaluation of Mixture of Expert Large Language and Vision Models
topic Machine Learning
Performance
url https://arxiv.org/abs/2508.17467