Relay Buffer Independent Communication over Pooled HBM for Efficient MoE Inference on Ascend
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hu, Tianlun, Hu, Tiancheng, Litang, Shengsheng, Wang, Sheng, Bao, Xiaoming, Li, Yuxing, Wang, Wei, Hu, Zhongzhe, Li, Lijun, Sun, Hongwei, Zhou, Jingbin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Remoe: Towards Efficient and Low-Cost MoE Inference in Serverless Computing
von: Liu, Wentao, et al.
Veröffentlicht: (2025)
von: Liu, Wentao, et al.
Veröffentlicht: (2025)
EC2MoE: Adaptive End-Cloud Pipeline Collaboration Enabling Scalable Mixture-of-Experts Inference
von: Yang, Zheming, et al.
Veröffentlicht: (2025)
von: Yang, Zheming, et al.
Veröffentlicht: (2025)
Accelerating Distributed MoE Training and Inference with Lina
von: Li, Jiamin, et al.
Veröffentlicht: (2022)
von: Li, Jiamin, et al.
Veröffentlicht: (2022)
MoE-Compression: How the Compression Error of Experts Affects the Inference Accuracy of MoE Model?
von: Ma, Songkai, et al.
Veröffentlicht: (2025)
von: Ma, Songkai, et al.
Veröffentlicht: (2025)
OD-MoE: On-Demand Expert Loading for Cacheless Edge-Distributed MoE Inference
von: Wang, Liujianfu, et al.
Veröffentlicht: (2025)
von: Wang, Liujianfu, et al.
Veröffentlicht: (2025)
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
von: Hu, Jianmin, et al.
Veröffentlicht: (2025)
von: Hu, Jianmin, et al.
Veröffentlicht: (2025)
EPS-MoE: Expert Pipeline Scheduler for Cost-Efficient MoE Inference
von: Qian, Yulei, et al.
Veröffentlicht: (2024)
von: Qian, Yulei, et al.
Veröffentlicht: (2024)
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs
von: Cao, Shiyi, et al.
Veröffentlicht: (2024)
von: Cao, Shiyi, et al.
Veröffentlicht: (2024)
SP-MoE: Speculative Decoding and Prefetching for Accelerating MoE-based Model Inference
von: Chen, Liangkun, et al.
Veröffentlicht: (2025)
von: Chen, Liangkun, et al.
Veröffentlicht: (2025)
Janus: Disaggregating Attention and Experts for Scalable MoE Inference
von: Zhang, Zhexiang, et al.
Veröffentlicht: (2025)
von: Zhang, Zhexiang, et al.
Veröffentlicht: (2025)
ReviveMoE: Fast Recovery for Hardware Failures in Large-Scale MoE LLM Inference Deployments
von: Li, Haley, et al.
Veröffentlicht: (2026)
von: Li, Haley, et al.
Veröffentlicht: (2026)
Staleness-Centric Optimizations for Parallel Diffusion MoE Inference
von: Luo, Jiajun, et al.
Veröffentlicht: (2024)
von: Luo, Jiajun, et al.
Veröffentlicht: (2024)
GRACE-MoE: Grouping and Replication with Locality-Aware Routing for Efficient Distributed MoE Inference
von: Han, Yu, et al.
Veröffentlicht: (2025)
von: Han, Yu, et al.
Veröffentlicht: (2025)
Surviving Partial Rank Failures in Wide Expert-Parallel MoE Inference
von: Sun, Xun, et al.
Veröffentlicht: (2026)
von: Sun, Xun, et al.
Veröffentlicht: (2026)
HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE Inference
von: Zhong, Shuzhang, et al.
Veröffentlicht: (2025)
von: Zhong, Shuzhang, et al.
Veröffentlicht: (2025)
HarMoEny: Efficient Multi-GPU Inference of MoE Models
von: Doucet, Zachary, et al.
Veröffentlicht: (2025)
von: Doucet, Zachary, et al.
Veröffentlicht: (2025)
ENEC: A Lossless AI Model Compression Method Enabling Fast Inference on Ascend NPUs
von: Yang, Jinwu, et al.
Veröffentlicht: (2026)
von: Yang, Jinwu, et al.
Veröffentlicht: (2026)
ReaLB: Real-Time Load Balancing for Multimodal MoE Inference
von: Wang, Yingping, et al.
Veröffentlicht: (2026)
von: Wang, Yingping, et al.
Veröffentlicht: (2026)
One Pool, Two Caches: Adaptive HBM Partitioning for Accelerating Generative Recommender Serving
von: Yu, Wenjun, et al.
Veröffentlicht: (2026)
von: Yu, Wenjun, et al.
Veröffentlicht: (2026)
LSH-MoE: Communication-efficient MoE Training via Locality-Sensitive Hashing
von: Nie, Xiaonan, et al.
Veröffentlicht: (2024)
von: Nie, Xiaonan, et al.
Veröffentlicht: (2024)
Efficient MoE Inference with Fine-Grained Scheduling of Disaggregated Expert Parallelism
von: Pan, Xinglin, et al.
Veröffentlicht: (2025)
von: Pan, Xinglin, et al.
Veröffentlicht: (2025)
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
von: Hu, Cunchen, et al.
Veröffentlicht: (2024)
von: Hu, Cunchen, et al.
Veröffentlicht: (2024)
HOBBIT: A Mixed Precision Expert Offloading System for Fast MoE Inference
von: Tang, Peng, et al.
Veröffentlicht: (2024)
von: Tang, Peng, et al.
Veröffentlicht: (2024)
MoE-Gen: High-Throughput MoE Inference on a Single GPU with Module-Based Batching
von: Xu, Tairan, et al.
Veröffentlicht: (2025)
von: Xu, Tairan, et al.
Veröffentlicht: (2025)
Efficient CPU-GPU Collaborative Inference for MoE-based LLMs on Memory-Limited Systems
von: Huang, En-Ming, et al.
Veröffentlicht: (2025)
von: Huang, En-Ming, et al.
Veröffentlicht: (2025)
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement
von: Wu, Tian, et al.
Veröffentlicht: (2025)
von: Wu, Tian, et al.
Veröffentlicht: (2025)
DALI: A Workload-Aware Offloading Framework for Efficient MoE Inference on Local PCs
von: Zhu, Zeyu, et al.
Veröffentlicht: (2026)
von: Zhu, Zeyu, et al.
Veröffentlicht: (2026)
Occult: Optimizing Collaborative Communication across Experts for Accelerated Parallel MoE Training and Inference
von: Luo, Shuqing, et al.
Veröffentlicht: (2025)
von: Luo, Shuqing, et al.
Veröffentlicht: (2025)
Hexa-MoE: Efficient and Heterogeneous-aware Training for Mixture-of-Experts
von: Luo, Shuqing, et al.
Veröffentlicht: (2024)
von: Luo, Shuqing, et al.
Veröffentlicht: (2024)
DuoServe-MoE: Dual-Phase Expert Prefetch and Caching for LLM Inference QoS Assurance
von: Zhang, Yuning, et al.
Veröffentlicht: (2025)
von: Zhang, Yuning, et al.
Veröffentlicht: (2025)
Semantic Parallelism: Redefining Efficient MoE Inference via Model-Data Co-Scheduling
von: Li, Yan, et al.
Veröffentlicht: (2025)
von: Li, Yan, et al.
Veröffentlicht: (2025)
Making MoE-based LLM Inference Resilient with Tarragon
von: Zhang, Songyu, et al.
Veröffentlicht: (2026)
von: Zhang, Songyu, et al.
Veröffentlicht: (2026)
EPD-Serve: A Flexible Multimodal EPD Disaggregation Inference Serving System On Ascend
von: Bai, Fan, et al.
Veröffentlicht: (2026)
von: Bai, Fan, et al.
Veröffentlicht: (2026)
When MoE Meets Blockchain: A Trustworthy Distributed Framework of Large Models
von: Zhu, Weihao, et al.
Veröffentlicht: (2025)
von: Zhu, Weihao, et al.
Veröffentlicht: (2025)
SCOOT: SLO-Oriented Performance Tuning for LLM Inference Engines
von: Cheng, Ke, et al.
Veröffentlicht: (2024)
von: Cheng, Ke, et al.
Veröffentlicht: (2024)
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems
von: Wu, Qi, et al.
Veröffentlicht: (2026)
von: Wu, Qi, et al.
Veröffentlicht: (2026)
Accelerating MoE Model Inference with Expert Sharding
von: Balmau, Oana, et al.
Veröffentlicht: (2025)
von: Balmau, Oana, et al.
Veröffentlicht: (2025)
Revealing the Challenges of Attention-FFN Disaggregation for Modern MoE Models and Hardware Systems
von: Liu, Guowei, et al.
Veröffentlicht: (2026)
von: Liu, Guowei, et al.
Veröffentlicht: (2026)
D$^{2}$MoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM Serving
von: Wang, Haodong, et al.
Veröffentlicht: (2025)
von: Wang, Haodong, et al.
Veröffentlicht: (2025)
EPIC: Efficient Position-Independent Caching for Serving Large Language Models
von: Hu, Junhao, et al.
Veröffentlicht: (2024)
von: Hu, Junhao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Remoe: Towards Efficient and Low-Cost MoE Inference in Serverless Computing
von: Liu, Wentao, et al.
Veröffentlicht: (2025) -
EC2MoE: Adaptive End-Cloud Pipeline Collaboration Enabling Scalable Mixture-of-Experts Inference
von: Yang, Zheming, et al.
Veröffentlicht: (2025) -
Accelerating Distributed MoE Training and Inference with Lina
von: Li, Jiamin, et al.
Veröffentlicht: (2022) -
MoE-Compression: How the Compression Error of Experts Affects the Inference Accuracy of MoE Model?
von: Ma, Songkai, et al.
Veröffentlicht: (2025) -
OD-MoE: On-Demand Expert Loading for Cacheless Edge-Distributed MoE Inference
von: Wang, Liujianfu, et al.
Veröffentlicht: (2025)