DALI: A Workload-Aware Offloading Framework for Efficient MoE Inference on Local PCs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhu, Zeyu, Li, Gang, Wang, Peisong, Mo, Zitao, Pei, Minnan, Song, Zhuoran, Liang, Xiaoyao, Cheng, Jian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
GRACE-MoE: Grouping and Replication with Locality-Aware Routing for Efficient Distributed MoE Inference
von: Han, Yu, et al.
Veröffentlicht: (2025)
von: Han, Yu, et al.
Veröffentlicht: (2025)
FastGL: A GPU-Efficient Framework for Accelerating Sampling-Based GNN Training at Large Scale
von: Zhu, Zeyu, et al.
Veröffentlicht: (2024)
von: Zhu, Zeyu, et al.
Veröffentlicht: (2024)
DAOP: Data-Aware Offloading and Predictive Pre-Calculation for Efficient MoE Inference
von: Zhang, Yujie, et al.
Veröffentlicht: (2024)
von: Zhang, Yujie, et al.
Veröffentlicht: (2024)
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
von: Hu, Jianmin, et al.
Veröffentlicht: (2025)
von: Hu, Jianmin, et al.
Veröffentlicht: (2025)
Priority-Aware Preemptive Scheduling for Mixed-Priority Workloads in MoE Inference
von: Siavashi, Mohammad, et al.
Veröffentlicht: (2025)
von: Siavashi, Mohammad, et al.
Veröffentlicht: (2025)
MoA-Off: Adaptive Heterogeneous Modality-Aware Offloading with Edge-Cloud Collaboration for Efficient Multimodal LLM Inference
von: Yang, Zheming, et al.
Veröffentlicht: (2025)
von: Yang, Zheming, et al.
Veröffentlicht: (2025)
HarMoEny: Efficient Multi-GPU Inference of MoE Models
von: Doucet, Zachary, et al.
Veröffentlicht: (2025)
von: Doucet, Zachary, et al.
Veröffentlicht: (2025)
SP-MoE: Speculative Decoding and Prefetching for Accelerating MoE-based Model Inference
von: Chen, Liangkun, et al.
Veröffentlicht: (2025)
von: Chen, Liangkun, et al.
Veröffentlicht: (2025)
OD-MoE: On-Demand Expert Loading for Cacheless Edge-Distributed MoE Inference
von: Wang, Liujianfu, et al.
Veröffentlicht: (2025)
von: Wang, Liujianfu, et al.
Veröffentlicht: (2025)
HOBBIT: A Mixed Precision Expert Offloading System for Fast MoE Inference
von: Tang, Peng, et al.
Veröffentlicht: (2024)
von: Tang, Peng, et al.
Veröffentlicht: (2024)
LSH-MoE: Communication-efficient MoE Training via Locality-Sensitive Hashing
von: Nie, Xiaonan, et al.
Veröffentlicht: (2024)
von: Nie, Xiaonan, et al.
Veröffentlicht: (2024)
EPS-MoE: Expert Pipeline Scheduler for Cost-Efficient MoE Inference
von: Qian, Yulei, et al.
Veröffentlicht: (2024)
von: Qian, Yulei, et al.
Veröffentlicht: (2024)
Accelerating Distributed MoE Training and Inference with Lina
von: Li, Jiamin, et al.
Veröffentlicht: (2022)
von: Li, Jiamin, et al.
Veröffentlicht: (2022)
Efficient CPU-GPU Collaborative Inference for MoE-based LLMs on Memory-Limited Systems
von: Huang, En-Ming, et al.
Veröffentlicht: (2025)
von: Huang, En-Ming, et al.
Veröffentlicht: (2025)
Staleness-Centric Optimizations for Parallel Diffusion MoE Inference
von: Luo, Jiajun, et al.
Veröffentlicht: (2024)
von: Luo, Jiajun, et al.
Veröffentlicht: (2024)
Janus: Disaggregating Attention and Experts for Scalable MoE Inference
von: Zhang, Zhexiang, et al.
Veröffentlicht: (2025)
von: Zhang, Zhexiang, et al.
Veröffentlicht: (2025)
MemFine: Memory-Aware Fine-Grained Scheduling for MoE Training
von: Zhao, Lu, et al.
Veröffentlicht: (2025)
von: Zhao, Lu, et al.
Veröffentlicht: (2025)
MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism
von: Zhang, Zheng, et al.
Veröffentlicht: (2025)
von: Zhang, Zheng, et al.
Veröffentlicht: (2025)
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems
von: Wu, Qi, et al.
Veröffentlicht: (2026)
von: Wu, Qi, et al.
Veröffentlicht: (2026)
Lynx: Enabling Efficient MoE Inference through Dynamic Batch-Aware Expert Selection
von: Gupta, Vima, et al.
Veröffentlicht: (2024)
von: Gupta, Vima, et al.
Veröffentlicht: (2024)
HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE Inference
von: Zhong, Shuzhang, et al.
Veröffentlicht: (2025)
von: Zhong, Shuzhang, et al.
Veröffentlicht: (2025)
ViBE: Co-Optimizing Workload Skew and Hardware Variability for MoE Serving
von: Go, Seokjin, et al.
Veröffentlicht: (2026)
von: Go, Seokjin, et al.
Veröffentlicht: (2026)
ReaLB: Real-Time Load Balancing for Multimodal MoE Inference
von: Wang, Yingping, et al.
Veröffentlicht: (2026)
von: Wang, Yingping, et al.
Veröffentlicht: (2026)
Surviving Partial Rank Failures in Wide Expert-Parallel MoE Inference
von: Sun, Xun, et al.
Veröffentlicht: (2026)
von: Sun, Xun, et al.
Veröffentlicht: (2026)
OCTOPINF: Workload-Aware Inference Serving for Edge Video Analytics
von: Nguyen, Thanh-Tung, et al.
Veröffentlicht: (2025)
von: Nguyen, Thanh-Tung, et al.
Veröffentlicht: (2025)
MSAO: Adaptive Modality Sparsity-Aware Offloading with Edge-Cloud Collaboration for Efficient Multimodal LLM Inference
von: Yang, Zheming, et al.
Veröffentlicht: (2026)
von: Yang, Zheming, et al.
Veröffentlicht: (2026)
ReviveMoE: Fast Recovery for Hardware Failures in Large-Scale MoE LLM Inference Deployments
von: Li, Haley, et al.
Veröffentlicht: (2026)
von: Li, Haley, et al.
Veröffentlicht: (2026)
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement
von: Wu, Tian, et al.
Veröffentlicht: (2025)
von: Wu, Tian, et al.
Veröffentlicht: (2025)
Hexa-MoE: Efficient and Heterogeneous-aware Training for Mixture-of-Experts
von: Luo, Shuqing, et al.
Veröffentlicht: (2024)
von: Luo, Shuqing, et al.
Veröffentlicht: (2024)
MoE-Compression: How the Compression Error of Experts Affects the Inference Accuracy of MoE Model?
von: Ma, Songkai, et al.
Veröffentlicht: (2025)
von: Ma, Songkai, et al.
Veröffentlicht: (2025)
Remoe: Towards Efficient and Low-Cost MoE Inference in Serverless Computing
von: Liu, Wentao, et al.
Veröffentlicht: (2025)
von: Liu, Wentao, et al.
Veröffentlicht: (2025)
Efficient MoE Inference with Fine-Grained Scheduling of Disaggregated Expert Parallelism
von: Pan, Xinglin, et al.
Veröffentlicht: (2025)
von: Pan, Xinglin, et al.
Veröffentlicht: (2025)
Revealing the Challenges of Attention-FFN Disaggregation for Modern MoE Models and Hardware Systems
von: Liu, Guowei, et al.
Veröffentlicht: (2026)
von: Liu, Guowei, et al.
Veröffentlicht: (2026)
Offloading Artificial Intelligence Workloads across the Computing Continuum by means of Active Storage Systems
von: Barceló, Alex, et al.
Veröffentlicht: (2025)
von: Barceló, Alex, et al.
Veröffentlicht: (2025)
Accelerating Mixture-of-Experts Inference by Hiding Offloading Latency with Speculative Decoding
von: Wang, Zhibin, et al.
Veröffentlicht: (2025)
von: Wang, Zhibin, et al.
Veröffentlicht: (2025)
MoE-Gen: High-Throughput MoE Inference on a Single GPU with Module-Based Batching
von: Xu, Tairan, et al.
Veröffentlicht: (2025)
von: Xu, Tairan, et al.
Veröffentlicht: (2025)
PROBE: Co-Balancing Computation and Communication in MoE Inference via Real-Time Predictive Prefetching
von: Zhu, Qianchao, et al.
Veröffentlicht: (2026)
von: Zhu, Qianchao, et al.
Veröffentlicht: (2026)
DuoServe-MoE: Dual-Phase Expert Prefetch and Caching for LLM Inference QoS Assurance
von: Zhang, Yuning, et al.
Veröffentlicht: (2025)
von: Zhang, Yuning, et al.
Veröffentlicht: (2025)
Static Batching of Irregular Workloads on GPUs: Framework and Application to Efficient MoE Model Inference
von: Li, Yinghan, et al.
Veröffentlicht: (2025)
von: Li, Yinghan, et al.
Veröffentlicht: (2025)
When MoE Meets Blockchain: A Trustworthy Distributed Framework of Large Models
von: Zhu, Weihao, et al.
Veröffentlicht: (2025)
von: Zhu, Weihao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
GRACE-MoE: Grouping and Replication with Locality-Aware Routing for Efficient Distributed MoE Inference
von: Han, Yu, et al.
Veröffentlicht: (2025) -
FastGL: A GPU-Efficient Framework for Accelerating Sampling-Based GNN Training at Large Scale
von: Zhu, Zeyu, et al.
Veröffentlicht: (2024) -
DAOP: Data-Aware Offloading and Predictive Pre-Calculation for Efficient MoE Inference
von: Zhang, Yujie, et al.
Veröffentlicht: (2024) -
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
von: Hu, Jianmin, et al.
Veröffentlicht: (2025) -
Priority-Aware Preemptive Scheduling for Mixed-Priority Workloads in MoE Inference
von: Siavashi, Mohammad, et al.
Veröffentlicht: (2025)