The $qs$ Inequality: Quantifying the Double Penalty of Mixture-of-Experts at Inference
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Adhinarayanan, Vignesh, Jayasena, Nuwan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Coordinated Reinforcement Learning Prefetching Architecture for Multicore Systems
von: Siddiqui, Mohammed Humaid, et al.
Veröffentlicht: (2025)
von: Siddiqui, Mohammed Humaid, et al.
Veröffentlicht: (2025)
Global Optimizations & Lightweight Dynamic Logic for Concurrency
von: Pati, Suchita, et al.
Veröffentlicht: (2024)
von: Pati, Suchita, et al.
Veröffentlicht: (2024)
T3: Transparent Tracking & Triggering for Fine-grained Overlap of Compute & Collectives
von: Pati, Suchita, et al.
Veröffentlicht: (2024)
von: Pati, Suchita, et al.
Veröffentlicht: (2024)
Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory
von: Jo, Myeong Jun
Veröffentlicht: (2026)
von: Jo, Myeong Jun
Veröffentlicht: (2026)
Cross-Platform Fused MoE Dispatch in Triton: Portable Expert Routing Without CUDA
von: Mitra, Subhadip
Veröffentlicht: (2026)
von: Mitra, Subhadip
Veröffentlicht: (2026)
SLED: A Speculative LLM Decoding Framework for Efficient Edge Serving
von: Li, Xiangchen, et al.
Veröffentlicht: (2025)
von: Li, Xiangchen, et al.
Veröffentlicht: (2025)
Deep Learning Model Deployment in Multiple Cloud Providers: an Exploratory Study Using Low Computing Power Environments
von: Lemos, Elayne, et al.
Veröffentlicht: (2025)
von: Lemos, Elayne, et al.
Veröffentlicht: (2025)
Splitwise: Collaborative Edge-Cloud Inference for LLMs via Lyapunov-Assisted DRL
von: Younesi, Abolfazl, et al.
Veröffentlicht: (2025)
von: Younesi, Abolfazl, et al.
Veröffentlicht: (2025)
Efficient and Scalable Architecture for Multiple-chip Implementation of Simulated Bifurcation Machines
von: Kashimata, Tomoya, et al.
Veröffentlicht: (2023)
von: Kashimata, Tomoya, et al.
Veröffentlicht: (2023)
Optimizing Multi-DNN Inference on Mobile Devices through Heterogeneous Processor Co-Execution
von: Gao, Yunquan, et al.
Veröffentlicht: (2025)
von: Gao, Yunquan, et al.
Veröffentlicht: (2025)
Accelerating State-Vector Quantum Simulation on Integrated GPUs via Cache Locality Optimization: A Cross-Architecture Evaluation
von: Thomaz, Gabriel Fernandes, et al.
Veröffentlicht: (2026)
von: Thomaz, Gabriel Fernandes, et al.
Veröffentlicht: (2026)
Predictive Multi-Tier Memory Management for KV Cache in Large-Scale GPU Inference
von: Ganjihal, Sanjeev Rao
Veröffentlicht: (2026)
von: Ganjihal, Sanjeev Rao
Veröffentlicht: (2026)
Accelerating Precise End-to-End Simulation: Latency-Sensitive Many-core System Modeling
von: Li, Yinrong, et al.
Veröffentlicht: (2026)
von: Li, Yinrong, et al.
Veröffentlicht: (2026)
Guess-Verify-Refine: Data-Aware Top-K for Sparse-Attention Decoding on Blackwell via Temporal Correlation
von: Cheng, Long, et al.
Veröffentlicht: (2026)
von: Cheng, Long, et al.
Veröffentlicht: (2026)
NCCL EP: Towards a Unified Expert Parallel Communication API for NCCL
von: Goldman, Amos, et al.
Veröffentlicht: (2026)
von: Goldman, Amos, et al.
Veröffentlicht: (2026)
Towards Greener LLMs: Bringing Energy-Efficiency to the Forefront of LLM Inference
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2024)
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2024)
FPGA-Accelerated Lock Management and Transaction Processing: Architecture, Optimization, and Design Space Exploration
von: Zhu, Shien, et al.
Veröffentlicht: (2026)
von: Zhu, Shien, et al.
Veröffentlicht: (2026)
FedPhD: Federated Pruning with Hierarchical Learning of Diffusion Models
von: Long, Qianyu, et al.
Veröffentlicht: (2025)
von: Long, Qianyu, et al.
Veröffentlicht: (2025)
Accelerating Frontier MoE Training with 3D Integrated Optics
von: Bernadskiy, Mikhail, et al.
Veröffentlicht: (2025)
von: Bernadskiy, Mikhail, et al.
Veröffentlicht: (2025)
SGDRC: Software-Defined Dynamic Resource Control for Concurrent DNN Inference on NVIDIA GPUs
von: Zhang, Yongkang, et al.
Veröffentlicht: (2024)
von: Zhang, Yongkang, et al.
Veröffentlicht: (2024)
Next-Generation Event-Driven Architectures: Performance, Scalability, and Intelligent Orchestration Across Messaging Frameworks
von: Arafat, Jahidul, et al.
Veröffentlicht: (2025)
von: Arafat, Jahidul, et al.
Veröffentlicht: (2025)
When Does Global Attention Help? A Unified Empirical Study on Atomistic Graph Learning
von: Chowdhury, Arindam, et al.
Veröffentlicht: (2025)
von: Chowdhury, Arindam, et al.
Veröffentlicht: (2025)
Optimizing Foundation Model Inference on a Many-tiny-core Open-source RISC-V Platform
von: Potocnik, Viviane, et al.
Veröffentlicht: (2024)
von: Potocnik, Viviane, et al.
Veröffentlicht: (2024)
Quantifying the Performance Gap for Simple Versus Optimal Dynamic Server Allocation Policies
von: Carlsson, Niklas, et al.
Veröffentlicht: (2025)
von: Carlsson, Niklas, et al.
Veröffentlicht: (2025)
Partitioned Neural Network Training via Synthetic Intermediate Labels
von: Karadağ, Cevat Volkan, et al.
Veröffentlicht: (2024)
von: Karadağ, Cevat Volkan, et al.
Veröffentlicht: (2024)
Architecture-Aware LLM Inference Optimization on AMD Instinct GPUs: A Comprehensive Benchmark and Deployment Study
von: Georgiou, Athos
Veröffentlicht: (2026)
von: Georgiou, Athos
Veröffentlicht: (2026)
Fusing Depthwise and Pointwise Convolutions for Efficient Inference on GPUs
von: Qararyah, Fareed, et al.
Veröffentlicht: (2024)
von: Qararyah, Fareed, et al.
Veröffentlicht: (2024)
Profiling Concurrent Vision Inference Workloads on NVIDIA Jetson -- Extended
von: Chakraborty, Abhinaba, et al.
Veröffentlicht: (2025)
von: Chakraborty, Abhinaba, et al.
Veröffentlicht: (2025)
RapidStream IR: Infrastructure for FPGA High-Level Physical Synthesis
von: Lau, Jason, et al.
Veröffentlicht: (2024)
von: Lau, Jason, et al.
Veröffentlicht: (2024)
Splitwise: Efficient generative LLM inference using phase splitting
von: Patel, Pratyush, et al.
Veröffentlicht: (2023)
von: Patel, Pratyush, et al.
Veröffentlicht: (2023)
Method for determining the acceleration of a parallel specialised computer system based on Amdahl's law
von: Filipchenko, Aleksandr S.
Veröffentlicht: (2024)
von: Filipchenko, Aleksandr S.
Veröffentlicht: (2024)
Evaluating Emerging AI/ML Accelerators: IPU, RDU, and NVIDIA/AMD GPUs
von: Peng, Hongwu, et al.
Veröffentlicht: (2023)
von: Peng, Hongwu, et al.
Veröffentlicht: (2023)
GPU-Initiated Networking for NCCL
von: Hamidouche, Khaled, et al.
Veröffentlicht: (2025)
von: Hamidouche, Khaled, et al.
Veröffentlicht: (2025)
GPT-OSS-20B: A Comprehensive Deployment-Centric Analysis of OpenAI's Open-Weight Mixture of Experts Model
von: Kumar, Deepak, et al.
Veröffentlicht: (2025)
von: Kumar, Deepak, et al.
Veröffentlicht: (2025)
Accelerating Geo-distributed Machine Learning with Network-Aware Adaptive Tree and Auxiliary Route
von: Li, Zonghang, et al.
Veröffentlicht: (2024)
von: Li, Zonghang, et al.
Veröffentlicht: (2024)
GigaAPI for GPU Parallelization
von: Suvarna, M., et al.
Veröffentlicht: (2025)
von: Suvarna, M., et al.
Veröffentlicht: (2025)
Performance Analysis of HPC applications on the Aurora Supercomputer: Exploring the Impact of HBM-Enabled Intel Xeon Max CPUs
von: Ibeid, Huda, et al.
Veröffentlicht: (2025)
von: Ibeid, Huda, et al.
Veröffentlicht: (2025)
UPMEM Unleashed: Software Secrets for Speed
von: Chmielewski, Krystian, et al.
Veröffentlicht: (2025)
von: Chmielewski, Krystian, et al.
Veröffentlicht: (2025)
Dissecting the NVIDIA Hopper Architecture through Microbenchmarking and Multiple Level Analysis
von: Luo, Weile, et al.
Veröffentlicht: (2025)
von: Luo, Weile, et al.
Veröffentlicht: (2025)
Experimental Assessment of Containers Running on Top of Virtual Machines
von: Aqasizade, Hossein, et al.
Veröffentlicht: (2024)
von: Aqasizade, Hossein, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Coordinated Reinforcement Learning Prefetching Architecture for Multicore Systems
von: Siddiqui, Mohammed Humaid, et al.
Veröffentlicht: (2025) -
Global Optimizations & Lightweight Dynamic Logic for Concurrency
von: Pati, Suchita, et al.
Veröffentlicht: (2024) -
T3: Transparent Tracking & Triggering for Fine-grained Overlap of Compute & Collectives
von: Pati, Suchita, et al.
Veröffentlicht: (2024) -
Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory
von: Jo, Myeong Jun
Veröffentlicht: (2026) -
Cross-Platform Fused MoE Dispatch in Triton: Portable Expert Routing Without CUDA
von: Mitra, Subhadip
Veröffentlicht: (2026)