MoE-Inference-Bench: Performance Evaluation of Mixture of Expert Large Language and Vision Models
Fuente:
arXiv
Saved in:
| Main Authors: | Chitty-Venkata, Krishna Teja, Howland, Sylvia, Azar, Golara, Soboleva, Daria, Vassilieva, Natalia, Raskar, Siddhisanket, Emani, Murali, Vishwanath, Venkatram |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LExI: Layer-Adaptive Active Experts for Efficient MoE Model Inference
by: Chitty-Venkata, Krishna Teja, et al.
Published: (2025)
by: Chitty-Venkata, Krishna Teja, et al.
Published: (2025)
PreLoRA: Hybrid Pre-training of Vision Transformers with Full Training and Low-Rank Adapters
by: Thapa, Krishu K, et al.
Published: (2025)
by: Thapa, Krishu K, et al.
Published: (2025)
LLM-Inference-Bench: Inference Benchmarking of Large Language Models on AI Accelerators
by: Chitty-Venkata, Krishna Teja, et al.
Published: (2024)
by: Chitty-Venkata, Krishna Teja, et al.
Published: (2024)
LangVision-LoRA-NAS: Neural Architecture Search for Variable LoRA Rank in Vision Language Models
by: Chitty-Venkata, Krishna Teja, et al.
Published: (2025)
by: Chitty-Venkata, Krishna Teja, et al.
Published: (2025)
MoPEQ: Mixture of Mixed Precision Quantized Experts
by: Chitty-Venkata, Krishna Teja, et al.
Published: (2025)
by: Chitty-Venkata, Krishna Teja, et al.
Published: (2025)
BaKlaVa -- Budgeted Allocation of KV cache for Long-context Inference
by: Gulhan, Ahmed Burak, et al.
Published: (2025)
by: Gulhan, Ahmed Burak, et al.
Published: (2025)
ImageNet-Think-250K: A Large-Scale Synthetic Dataset for Multimodal Reasoning for Vision Language Models
by: Chitty-Venkata, Krishna Teja, et al.
Published: (2025)
by: Chitty-Venkata, Krishna Teja, et al.
Published: (2025)
Swimba: Switch Mamba Model Scales State Space Models
by: Du, Zhixu, et al.
Published: (2026)
by: Du, Zhixu, et al.
Published: (2026)
PagedEviction: Structured Block-wise KV Cache Pruning for Efficient Large Language Model Inference
by: Chitty-Venkata, Krishna Teja, et al.
Published: (2025)
by: Chitty-Venkata, Krishna Teja, et al.
Published: (2025)
MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache
by: Xue, Leyang, et al.
Published: (2024)
by: Xue, Leyang, et al.
Published: (2024)
HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing
by: Huang, Haochen, et al.
Published: (2025)
by: Huang, Haochen, et al.
Published: (2025)
ExpertFlow: Adaptive Expert Scheduling and Memory Coordination for Efficient MoE Inference
by: Shen, Zixu, et al.
Published: (2025)
by: Shen, Zixu, et al.
Published: (2025)
Dynamic Expert Quantization for Scalable Mixture-of-Experts Inference
by: Chu, Kexin, et al.
Published: (2025)
by: Chu, Kexin, et al.
Published: (2025)
MoEITS: A Green AI approach for simplifying MoE-LLMs
by: Balderas, Luis, et al.
Published: (2026)
by: Balderas, Luis, et al.
Published: (2026)
Cross-Platform Fused MoE Dispatch in Triton: Portable Expert Routing Without CUDA
by: Mitra, Subhadip
Published: (2026)
by: Mitra, Subhadip
Published: (2026)
Mixture of Experts with Mixture of Precisions for Tuning Quality of Service
by: Imani, HamidReza, et al.
Published: (2024)
by: Imani, HamidReza, et al.
Published: (2024)
Performance Characterization of Expert Router for Scalable LLM Inference
by: Pichlmeier, Josef, et al.
Published: (2024)
by: Pichlmeier, Josef, et al.
Published: (2024)
Fast NF4 Dequantization Kernels for Large Language Model Inference
by: Qi, Xiangbo, et al.
Published: (2026)
by: Qi, Xiangbo, et al.
Published: (2026)
KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation
by: Jiang, Chaoyi, et al.
Published: (2024)
by: Jiang, Chaoyi, et al.
Published: (2024)
WritePolicyBench: Benchmarking Memory Write Policies under Byte Budgets
by: Cham, Edgard El
Published: (2026)
by: Cham, Edgard El
Published: (2026)
Network Anatomy and Real-Time Measurement of Nvidia GeForce NOW Cloud Gaming
by: Lyu, Minzhao, et al.
Published: (2024)
by: Lyu, Minzhao, et al.
Published: (2024)
Assessing Home-Field Advantage in the Presidents Cup: Impact on Competitive Balance and Team Performance
by: Ehrlich, Justin, et al.
Published: (2025)
by: Ehrlich, Justin, et al.
Published: (2025)
BranchBench: Aligning Database Branching with Agentic Demands
by: Ang, Elaine, et al.
Published: (2026)
by: Ang, Elaine, et al.
Published: (2026)
Bench360: Benchmarking Local LLM Inference from 360 Degrees
by: Stuhlmann, Linus, et al.
Published: (2025)
by: Stuhlmann, Linus, et al.
Published: (2025)
Dissecting Embedding Bag Performance in DLRM Inference
by: Ambati, Chandrish, et al.
Published: (2025)
by: Ambati, Chandrish, et al.
Published: (2025)
Statistical Modeling and Uncertainty Estimation of LLM Inference Systems
by: Ray, Kaustabha, et al.
Published: (2025)
by: Ray, Kaustabha, et al.
Published: (2025)
Machine Learning-Guided Memory Optimization for DLRM Inference on Tiered Memory
by: Ren, Jie, et al.
Published: (2025)
by: Ren, Jie, et al.
Published: (2025)
Characterize LSM-tree Compaction Performance via On-Device LLM Inference
by: Ding, Jiabiao, et al.
Published: (2026)
by: Ding, Jiabiao, et al.
Published: (2026)
Meta-Metrics and Best Practices for System-Level Inference Performance Benchmarking
by: Salaria, Shweta, et al.
Published: (2025)
by: Salaria, Shweta, et al.
Published: (2025)
The $qs$ Inequality: Quantifying the Double Penalty of Mixture-of-Experts at Inference
by: Adhinarayanan, Vignesh, et al.
Published: (2026)
by: Adhinarayanan, Vignesh, et al.
Published: (2026)
Explainable Port Mapping Inference with Sparse Performance Counters for AMD's Zen Architectures
by: Ritter, Fabian, et al.
Published: (2024)
by: Ritter, Fabian, et al.
Published: (2024)
Profiling Large Language Model Inference on Apple Silicon: A Quantization Perspective
by: Benazir, Afsara, et al.
Published: (2025)
by: Benazir, Afsara, et al.
Published: (2025)
SparseInfer: Training-free Prediction of Activation Sparsity for Fast LLM Inference
by: Shin, Jiho, et al.
Published: (2024)
by: Shin, Jiho, et al.
Published: (2024)
CoServe: Efficient Collaboration-of-Experts (CoE) Model Inference with Limited Memory
by: Suo, Jiashun, et al.
Published: (2025)
by: Suo, Jiashun, et al.
Published: (2025)
Inference performance evaluation for LLMs on edge devices with a novel benchmarking framework and metric
by: Chen, Hao, et al.
Published: (2025)
by: Chen, Hao, et al.
Published: (2025)
Elastic MoE: Unlocking the Inference-Time Scalability of Mixture-of-Experts
by: Gu, Naibin, et al.
Published: (2025)
by: Gu, Naibin, et al.
Published: (2025)
H2EAL: Hybrid-Bonding Architecture with Hybrid Sparse Attention for Efficient Long-Context LLM Inference
by: Fu, Zizhuo, et al.
Published: (2025)
by: Fu, Zizhuo, et al.
Published: (2025)
KernelBenchX: A Comprehensive Benchmark for Evaluating LLM-Generated GPU Kernels
by: Wang, Han, et al.
Published: (2026)
by: Wang, Han, et al.
Published: (2026)
HI-MoE: Hierarchical Instance-Conditioned Mixture-of-Experts for Object Detection
by: Vashkelis, Vadim, et al.
Published: (2026)
by: Vashkelis, Vadim, et al.
Published: (2026)
Iterative Layer Pruning for Efficient Translation Inference
by: Moslem, Yasmin, et al.
Published: (2025)
by: Moslem, Yasmin, et al.
Published: (2025)
Similar Items
-
LExI: Layer-Adaptive Active Experts for Efficient MoE Model Inference
by: Chitty-Venkata, Krishna Teja, et al.
Published: (2025) -
PreLoRA: Hybrid Pre-training of Vision Transformers with Full Training and Low-Rank Adapters
by: Thapa, Krishu K, et al.
Published: (2025) -
LLM-Inference-Bench: Inference Benchmarking of Large Language Models on AI Accelerators
by: Chitty-Venkata, Krishna Teja, et al.
Published: (2024) -
LangVision-LoRA-NAS: Neural Architecture Search for Variable LoRA Rank in Vision Language Models
by: Chitty-Venkata, Krishna Teja, et al.
Published: (2025) -
MoPEQ: Mixture of Mixed Precision Quantized Experts
by: Chitty-Venkata, Krishna Teja, et al.
Published: (2025)