Priority-Aware Preemptive Scheduling for Mixed-Priority Workloads in MoE Inference
Fuente:
arXiv
Salvato in:
| Autori principali: | Siavashi, Mohammad, Dindarloo, Faezeh Keshmiri, Kostic, Dejan, Chiesa, Marco |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Blink: CPU-Free LLM Inference by Delegating the Serving Stack to GPU and SmartNIC
di: Siavashi, Mohammad, et al.
Pubblicazione: (2026)
di: Siavashi, Mohammad, et al.
Pubblicazione: (2026)
DALI: A Workload-Aware Offloading Framework for Efficient MoE Inference on Local PCs
di: Zhu, Zeyu, et al.
Pubblicazione: (2026)
di: Zhu, Zeyu, et al.
Pubblicazione: (2026)
GCAPS: GPU Context-Aware Preemptive Priority-based Scheduling for Real-Time Tasks
di: Wang, Yidi, et al.
Pubblicazione: (2024)
di: Wang, Yidi, et al.
Pubblicazione: (2024)
Priority-Aware Model-Distributed Inference at Edge Networks
di: Li, Teng, et al.
Pubblicazione: (2024)
di: Li, Teng, et al.
Pubblicazione: (2024)
HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE Inference
di: Zhong, Shuzhang, et al.
Pubblicazione: (2025)
di: Zhong, Shuzhang, et al.
Pubblicazione: (2025)
PecSched: Preemptive and Efficient Cluster Scheduling for LLM Inference
di: Zhang, Zeyu, et al.
Pubblicazione: (2024)
di: Zhang, Zeyu, et al.
Pubblicazione: (2024)
ViBE: Co-Optimizing Workload Skew and Hardware Variability for MoE Serving
di: Go, Seokjin, et al.
Pubblicazione: (2026)
di: Go, Seokjin, et al.
Pubblicazione: (2026)
HOBBIT: A Mixed Precision Expert Offloading System for Fast MoE Inference
di: Tang, Peng, et al.
Pubblicazione: (2024)
di: Tang, Peng, et al.
Pubblicazione: (2024)
MoE-Compression: How the Compression Error of Experts Affects the Inference Accuracy of MoE Model?
di: Ma, Songkai, et al.
Pubblicazione: (2025)
di: Ma, Songkai, et al.
Pubblicazione: (2025)
Unleashing the Power of Preemptive Priority-based Scheduling for Real-Time GPU Tasks
di: Wang, Yidi, et al.
Pubblicazione: (2024)
di: Wang, Yidi, et al.
Pubblicazione: (2024)
MoE-Gen: High-Throughput MoE Inference on a Single GPU with Module-Based Batching
di: Xu, Tairan, et al.
Pubblicazione: (2025)
di: Xu, Tairan, et al.
Pubblicazione: (2025)
DAOP: Data-Aware Offloading and Predictive Pre-Calculation for Efficient MoE Inference
di: Zhang, Yujie, et al.
Pubblicazione: (2024)
di: Zhang, Yujie, et al.
Pubblicazione: (2024)
Optimizing Performance on Trinity Utilizing Machine Learning, Proxy Applications and Scheduling Priorities
di: Romero, Phil
Pubblicazione: (2024)
di: Romero, Phil
Pubblicazione: (2024)
Lynx: Enabling Efficient MoE Inference through Dynamic Batch-Aware Expert Selection
di: Gupta, Vima, et al.
Pubblicazione: (2024)
di: Gupta, Vima, et al.
Pubblicazione: (2024)
Making MoE-based LLM Inference Resilient with Tarragon
di: Zhang, Songyu, et al.
Pubblicazione: (2026)
di: Zhang, Songyu, et al.
Pubblicazione: (2026)
FIKIT: Priority-Based Real-time GPU Multi-tasking Scheduling with Kernel Identification
di: Wu, Wenqing
Pubblicazione: (2023)
di: Wu, Wenqing
Pubblicazione: (2023)
From Tokens to Layers: Redefining Stall-Free Scheduling for MoE Serving with Layered Prefill
di: Lee, Gunjun, et al.
Pubblicazione: (2025)
di: Lee, Gunjun, et al.
Pubblicazione: (2025)
eMoE: Task-aware Memory Efficient Mixture-of-Experts-Based (MoE) Model Inference
di: Tairin, Suraiya, et al.
Pubblicazione: (2025)
di: Tairin, Suraiya, et al.
Pubblicazione: (2025)
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs
di: Cao, Shiyi, et al.
Pubblicazione: (2024)
di: Cao, Shiyi, et al.
Pubblicazione: (2024)
MoE Parallel Folding: Heterogeneous Parallelism Mappings for Efficient Large-Scale MoE Model Training with Megatron Core
di: Liu, Dennis, et al.
Pubblicazione: (2025)
di: Liu, Dennis, et al.
Pubblicazione: (2025)
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
di: Hu, Jianmin, et al.
Pubblicazione: (2025)
di: Hu, Jianmin, et al.
Pubblicazione: (2025)
Semantic Parallelism: Redefining Efficient MoE Inference via Model-Data Co-Scheduling
di: Li, Yan, et al.
Pubblicazione: (2025)
di: Li, Yan, et al.
Pubblicazione: (2025)
Occult: Optimizing Collaborative Communication across Experts for Accelerated Parallel MoE Training and Inference
di: Luo, Shuqing, et al.
Pubblicazione: (2025)
di: Luo, Shuqing, et al.
Pubblicazione: (2025)
Relay Buffer Independent Communication over Pooled HBM for Efficient MoE Inference on Ascend
di: Hu, Tianlun, et al.
Pubblicazione: (2026)
di: Hu, Tianlun, et al.
Pubblicazione: (2026)
EPS-MoE: Expert Pipeline Scheduler for Cost-Efficient MoE Inference
di: Qian, Yulei, et al.
Pubblicazione: (2024)
di: Qian, Yulei, et al.
Pubblicazione: (2024)
Metronome: Efficient Scheduling for Periodic Traffic Jobs with Network and Priority Awareness
di: Jiang, Hao, et al.
Pubblicazione: (2025)
di: Jiang, Hao, et al.
Pubblicazione: (2025)
Accelerating MoE Model Inference with Expert Sharding
di: Balmau, Oana, et al.
Pubblicazione: (2025)
di: Balmau, Oana, et al.
Pubblicazione: (2025)
GRACE-MoE: Grouping and Replication with Locality-Aware Routing for Efficient Distributed MoE Inference
di: Han, Yu, et al.
Pubblicazione: (2025)
di: Han, Yu, et al.
Pubblicazione: (2025)
HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap
di: Lin, Wenxiang, et al.
Pubblicazione: (2025)
di: Lin, Wenxiang, et al.
Pubblicazione: (2025)
ZipMoE: Efficient On-Device MoE Serving via Lossless Compression and Cache-Affinity Scheduling
di: Yang, Yuchen, et al.
Pubblicazione: (2026)
di: Yang, Yuchen, et al.
Pubblicazione: (2026)
SiDA-MoE: Sparsity-Inspired Data-Aware Serving for Efficient and Scalable Large Mixture-of-Experts Models
di: Du, Zhixu, et al.
Pubblicazione: (2023)
di: Du, Zhixu, et al.
Pubblicazione: (2023)
Preemption Aware Task Scheduling for Priority and Deadline Constrained DNN Inference Task Offloading in Homogeneous Mobile-Edge Networks
di: Cotter, Jamie, et al.
Pubblicazione: (2025)
di: Cotter, Jamie, et al.
Pubblicazione: (2025)
CARIn: Constraint-Aware and Responsive Inference on Heterogeneous Devices for Single- and Multi-DNN Workloads
di: Panopoulos, Ioannis, et al.
Pubblicazione: (2024)
di: Panopoulos, Ioannis, et al.
Pubblicazione: (2024)
MemFine: Memory-Aware Fine-Grained Scheduling for MoE Training
di: Zhao, Lu, et al.
Pubblicazione: (2025)
di: Zhao, Lu, et al.
Pubblicazione: (2025)
RankMap: Priority-Aware Multi-DNN Manager for Heterogeneous Embedded Devices
di: Karatzas, Andreas, et al.
Pubblicazione: (2024)
di: Karatzas, Andreas, et al.
Pubblicazione: (2024)
Topology-aware Preemptive Scheduling for Co-located LLM Workloads
di: Zhang, Ping, et al.
Pubblicazione: (2024)
di: Zhang, Ping, et al.
Pubblicazione: (2024)
MoE-CAP: Benchmarking Cost, Accuracy and Performance of Sparse Mixture-of-Experts Systems
di: Jiang, Yinsicheng, et al.
Pubblicazione: (2025)
di: Jiang, Yinsicheng, et al.
Pubblicazione: (2025)
MoE-CAP: Benchmarking Cost, Accuracy and Performance of Sparse Mixture-of-Experts Systems
di: Jiang, Yinsicheng, et al.
Pubblicazione: (2024)
di: Jiang, Yinsicheng, et al.
Pubblicazione: (2024)
ExeGPT: Constraint-Aware Resource Scheduling for LLM Inference
di: Oh, Hyungjun, et al.
Pubblicazione: (2024)
di: Oh, Hyungjun, et al.
Pubblicazione: (2024)
PROSERVE: Unified Multi-Priority Request Scheduling for LLM Serving
di: Huang, Weizhe, et al.
Pubblicazione: (2025)
di: Huang, Weizhe, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Blink: CPU-Free LLM Inference by Delegating the Serving Stack to GPU and SmartNIC
di: Siavashi, Mohammad, et al.
Pubblicazione: (2026) -
DALI: A Workload-Aware Offloading Framework for Efficient MoE Inference on Local PCs
di: Zhu, Zeyu, et al.
Pubblicazione: (2026) -
GCAPS: GPU Context-Aware Preemptive Priority-based Scheduling for Real-Time Tasks
di: Wang, Yidi, et al.
Pubblicazione: (2024) -
Priority-Aware Model-Distributed Inference at Edge Networks
di: Li, Teng, et al.
Pubblicazione: (2024) -
HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE Inference
di: Zhong, Shuzhang, et al.
Pubblicazione: (2025)