Revealing the Challenges of Attention-FFN Disaggregation for Modern MoE Models and Hardware Systems
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Guowei, Li, Hongming, Guo, Yaning, Lyu, Yongxi, Zhou, Mo, Liu, Yi, Li, Zhaogeng, Wang, Yanpeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Janus: Disaggregating Attention and Experts for Scalable MoE Inference
von: Zhang, Zhexiang, et al.
Veröffentlicht: (2025)
von: Zhang, Zhexiang, et al.
Veröffentlicht: (2025)
How Far Can Disaggregation Go? A Design-Space Exploration of Attention-FFN Disaggregation for Efficient MoE LLM Serving
von: Wu, Hanjiang, et al.
Veröffentlicht: (2026)
von: Wu, Hanjiang, et al.
Veröffentlicht: (2026)
LSH-MoE: Communication-efficient MoE Training via Locality-Sensitive Hashing
von: Nie, Xiaonan, et al.
Veröffentlicht: (2024)
von: Nie, Xiaonan, et al.
Veröffentlicht: (2024)
ReviveMoE: Fast Recovery for Hardware Failures in Large-Scale MoE LLM Inference Deployments
von: Li, Haley, et al.
Veröffentlicht: (2026)
von: Li, Haley, et al.
Veröffentlicht: (2026)
MoE-Hub: Taming Software Complexity for Seamless MoE Overlap with Hardware-Accelerated Communication on Multi-GPU Systems
von: Zhou, Zhuoshan, et al.
Veröffentlicht: (2026)
von: Zhou, Zhuoshan, et al.
Veröffentlicht: (2026)
OD-MoE: On-Demand Expert Loading for Cacheless Edge-Distributed MoE Inference
von: Wang, Liujianfu, et al.
Veröffentlicht: (2025)
von: Wang, Liujianfu, et al.
Veröffentlicht: (2025)
Efficient CPU-GPU Collaborative Inference for MoE-based LLMs on Memory-Limited Systems
von: Huang, En-Ming, et al.
Veröffentlicht: (2025)
von: Huang, En-Ming, et al.
Veröffentlicht: (2025)
Efficient MoE Inference with Fine-Grained Scheduling of Disaggregated Expert Parallelism
von: Pan, Xinglin, et al.
Veröffentlicht: (2025)
von: Pan, Xinglin, et al.
Veröffentlicht: (2025)
Accelerating Distributed MoE Training and Inference with Lina
von: Li, Jiamin, et al.
Veröffentlicht: (2022)
von: Li, Jiamin, et al.
Veröffentlicht: (2022)
SP-MoE: Speculative Decoding and Prefetching for Accelerating MoE-based Model Inference
von: Chen, Liangkun, et al.
Veröffentlicht: (2025)
von: Chen, Liangkun, et al.
Veröffentlicht: (2025)
MoE-Lens: Towards the Hardware Limit of High-Throughput MoE LLM Serving Under Resource Constraints
von: Yuan, Yichao, et al.
Veröffentlicht: (2025)
von: Yuan, Yichao, et al.
Veröffentlicht: (2025)
ReaLB: Real-Time Load Balancing for Multimodal MoE Inference
von: Wang, Yingping, et al.
Veröffentlicht: (2026)
von: Wang, Yingping, et al.
Veröffentlicht: (2026)
GRACE-MoE: Grouping and Replication with Locality-Aware Routing for Efficient Distributed MoE Inference
von: Han, Yu, et al.
Veröffentlicht: (2025)
von: Han, Yu, et al.
Veröffentlicht: (2025)
INDIGO: Page Migration for Hardware Memory Disaggregation Across a Network
von: Patke, Archit, et al.
Veröffentlicht: (2025)
von: Patke, Archit, et al.
Veröffentlicht: (2025)
HOBBIT: A Mixed Precision Expert Offloading System for Fast MoE Inference
von: Tang, Peng, et al.
Veröffentlicht: (2024)
von: Tang, Peng, et al.
Veröffentlicht: (2024)
Hexa-MoE: Efficient and Heterogeneous-aware Training for Mixture-of-Experts
von: Luo, Shuqing, et al.
Veröffentlicht: (2024)
von: Luo, Shuqing, et al.
Veröffentlicht: (2024)
From Attention to Disaggregation: Tracing the Evolution of LLM Inference
von: Kumar, Madabattula Rajesh, et al.
Veröffentlicht: (2025)
von: Kumar, Madabattula Rajesh, et al.
Veröffentlicht: (2025)
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models
von: Wang, Wei, et al.
Veröffentlicht: (2024)
von: Wang, Wei, et al.
Veröffentlicht: (2024)
MoEBlaze: Breaking the Memory Wall for Efficient MoE Training on Modern GPUs
von: Zhang, Jiyuan, et al.
Veröffentlicht: (2026)
von: Zhang, Jiyuan, et al.
Veröffentlicht: (2026)
ViBE: Co-Optimizing Workload Skew and Hardware Variability for MoE Serving
von: Go, Seokjin, et al.
Veröffentlicht: (2026)
von: Go, Seokjin, et al.
Veröffentlicht: (2026)
FEPLB: Exploiting Copy Engines for Nearly Free MoE Load Balancing in Distributed Training
von: Qi, Shuyao, et al.
Veröffentlicht: (2026)
von: Qi, Shuyao, et al.
Veröffentlicht: (2026)
DisagMoE: Computation-Communication overlapped MoE Training via Disaggregated AF-Pipe Parallelism
von: Zeng, Zhichen, et al.
Veröffentlicht: (2026)
von: Zeng, Zhichen, et al.
Veröffentlicht: (2026)
EPS-MoE: Expert Pipeline Scheduler for Cost-Efficient MoE Inference
von: Qian, Yulei, et al.
Veröffentlicht: (2024)
von: Qian, Yulei, et al.
Veröffentlicht: (2024)
Proceedings of 3rd Workshop on Heterogeneous Composable and Disaggregated Systems
von: Pinto, Christian, et al.
Veröffentlicht: (2024)
von: Pinto, Christian, et al.
Veröffentlicht: (2024)
HarMoEny: Efficient Multi-GPU Inference of MoE Models
von: Doucet, Zachary, et al.
Veröffentlicht: (2025)
von: Doucet, Zachary, et al.
Veröffentlicht: (2025)
UniEP: Unified Expert-Parallel MoE MegaKernel for LLM Training
von: Zheng, Size, et al.
Veröffentlicht: (2026)
von: Zheng, Size, et al.
Veröffentlicht: (2026)
Sparse Checkpointing for Fast and Reliable MoE Training
von: Gandhi, Swapnil, et al.
Veröffentlicht: (2024)
von: Gandhi, Swapnil, et al.
Veröffentlicht: (2024)
vLLM-Omni: Fully Disaggregated Serving for Any-to-Any Multimodal Models
von: Yin, Peiqi, et al.
Veröffentlicht: (2026)
von: Yin, Peiqi, et al.
Veröffentlicht: (2026)
Lotus: Optimizing Disaggregated Transactions with Disaggregated Locks
von: Hu, Zhisheng, et al.
Veröffentlicht: (2025)
von: Hu, Zhisheng, et al.
Veröffentlicht: (2025)
MoE-Compression: How the Compression Error of Experts Affects the Inference Accuracy of MoE Model?
von: Ma, Songkai, et al.
Veröffentlicht: (2025)
von: Ma, Songkai, et al.
Veröffentlicht: (2025)
When MoE Meets Blockchain: A Trustworthy Distributed Framework of Large Models
von: Zhu, Weihao, et al.
Veröffentlicht: (2025)
von: Zhu, Weihao, et al.
Veröffentlicht: (2025)
HydraInfer: Hybrid Disaggregated Scheduling for Multimodal Large Language Model Serving
von: Dong, Xianzhe, et al.
Veröffentlicht: (2025)
von: Dong, Xianzhe, et al.
Veröffentlicht: (2025)
PROBE: Co-Balancing Computation and Communication in MoE Inference via Real-Time Predictive Prefetching
von: Zhu, Qianchao, et al.
Veröffentlicht: (2026)
von: Zhu, Qianchao, et al.
Veröffentlicht: (2026)
Fine-grained MoE Load Balancing with Linear Programming
von: Zhao, Chenqi, et al.
Veröffentlicht: (2025)
von: Zhao, Chenqi, et al.
Veröffentlicht: (2025)
Multi-Layer Scheduling for MoE-Based LLM Reasoning
von: Sun, Yifan, et al.
Veröffentlicht: (2026)
von: Sun, Yifan, et al.
Veröffentlicht: (2026)
Staleness-Centric Optimizations for Parallel Diffusion MoE Inference
von: Luo, Jiajun, et al.
Veröffentlicht: (2024)
von: Luo, Jiajun, et al.
Veröffentlicht: (2024)
Surviving Partial Rank Failures in Wide Expert-Parallel MoE Inference
von: Sun, Xun, et al.
Veröffentlicht: (2026)
von: Sun, Xun, et al.
Veröffentlicht: (2026)
MoE Parallel Folding: Heterogeneous Parallelism Mappings for Efficient Large-Scale MoE Model Training with Megatron Core
von: Liu, Dennis, et al.
Veröffentlicht: (2025)
von: Liu, Dennis, et al.
Veröffentlicht: (2025)
Expert-as-a-Service: Towards Efficient, Scalable, and Robust Large-scale MoE Serving
von: Liu, Ziming, et al.
Veröffentlicht: (2025)
von: Liu, Ziming, et al.
Veröffentlicht: (2025)
D$^{2}$MoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM Serving
von: Wang, Haodong, et al.
Veröffentlicht: (2025)
von: Wang, Haodong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Janus: Disaggregating Attention and Experts for Scalable MoE Inference
von: Zhang, Zhexiang, et al.
Veröffentlicht: (2025) -
How Far Can Disaggregation Go? A Design-Space Exploration of Attention-FFN Disaggregation for Efficient MoE LLM Serving
von: Wu, Hanjiang, et al.
Veröffentlicht: (2026) -
LSH-MoE: Communication-efficient MoE Training via Locality-Sensitive Hashing
von: Nie, Xiaonan, et al.
Veröffentlicht: (2024) -
ReviveMoE: Fast Recovery for Hardware Failures in Large-Scale MoE LLM Inference Deployments
von: Li, Haley, et al.
Veröffentlicht: (2026) -
MoE-Hub: Taming Software Complexity for Seamless MoE Overlap with Hardware-Accelerated Communication on Multi-GPU Systems
von: Zhou, Zhuoshan, et al.
Veröffentlicht: (2026)