Pangu Ultra MoE: How to Train Your Big MoE on Ascend NPUs
Fuente:
arXiv
Guardado en:
Ejemplares similares
Pangu Pro MoE: Mixture of Grouped Experts for Efficient Sparsity
por: Tang, Yehui, et al.
Publicado: (2025)
por: Tang, Yehui, et al.
Publicado: (2025)
Pangu Ultra: Pushing the Limits of Dense Large Language Models on Ascend NPUs
por: Yin, Yichun, et al.
Publicado: (2025)
por: Yin, Yichun, et al.
Publicado: (2025)
MoE-Prism: Disentangling Monolithic Experts for Elastic MoE Services via Model-System Co-Designs
por: Xia, Xinfeng, et al.
Publicado: (2025)
por: Xia, Xinfeng, et al.
Publicado: (2025)
Alloc-MoE: Budget-Aware Expert Activation Allocation for Efficient Mixture-of-Experts Inference
por: Liu, Baihui, et al.
Publicado: (2026)
por: Liu, Baihui, et al.
Publicado: (2026)
LLaDA-MoE: A Sparse MoE Diffusion Language Model
por: Zhu, Fengqi, et al.
Publicado: (2025)
por: Zhu, Fengqi, et al.
Publicado: (2025)
Training Report of TeleChat3-MoE
por: Liu, Xinzhang, et al.
Publicado: (2025)
por: Liu, Xinzhang, et al.
Publicado: (2025)
Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE
por: Huang, Haiduo, et al.
Publicado: (2025)
por: Huang, Haiduo, et al.
Publicado: (2025)
MoE-Compression: How the Compression Error of Experts Affects the Inference Accuracy of MoE Model?
por: Ma, Songkai, et al.
Publicado: (2025)
por: Ma, Songkai, et al.
Publicado: (2025)
Grove MoE: Towards Efficient and Superior MoE LLMs with Adjugate Experts
por: Wu, Haoyuan, et al.
Publicado: (2025)
por: Wu, Haoyuan, et al.
Publicado: (2025)
MoEless: Efficient MoE LLM Serving via Serverless Computing
por: Yu, Hanfei, et al.
Publicado: (2026)
por: Yu, Hanfei, et al.
Publicado: (2026)
Expert Divergence Learning for MoE-based Language Models
por: Li, Jiaang, et al.
Publicado: (2026)
por: Li, Jiaang, et al.
Publicado: (2026)
LSH-MoE: Communication-efficient MoE Training via Locality-Sensitive Hashing
por: Nie, Xiaonan, et al.
Publicado: (2024)
por: Nie, Xiaonan, et al.
Publicado: (2024)
LongScape: Advancing Long-Horizon Embodied World Models with Context-Aware MoE
por: Shang, Yu, et al.
Publicado: (2025)
por: Shang, Yu, et al.
Publicado: (2025)
Analytical FFN-to-MoE Restructuring via Activation Pattern Analysis
por: Pei, Zehua, et al.
Publicado: (2025)
por: Pei, Zehua, et al.
Publicado: (2025)
CRAM: Centroid-Routing and Adaptive MoE for Multimodal Continual Instruction Tuning
por: Tang, Jun-Tao, et al.
Publicado: (2026)
por: Tang, Jun-Tao, et al.
Publicado: (2026)
Static Batching of Irregular Workloads on GPUs: Framework and Application to Efficient MoE Model Inference
por: Li, Yinghan, et al.
Publicado: (2025)
por: Li, Yinghan, et al.
Publicado: (2025)
Unveiling Language Routing Isolation in Multilingual MoE Models for Interpretable Subnetwork Adaptation
por: Zheng, Kening, et al.
Publicado: (2026)
por: Zheng, Kening, et al.
Publicado: (2026)
MoESD: Unveil Speculative Decoding's Potential for Accelerating Sparse MoE
por: Huang, Zongle, et al.
Publicado: (2025)
por: Huang, Zongle, et al.
Publicado: (2025)
EAQuant: Enhancing Post-Training Quantization for MoE Models via Expert-Aware Optimization
por: Fu, Zhongqian, et al.
Publicado: (2025)
por: Fu, Zhongqian, et al.
Publicado: (2025)
Revealing the Challenges of Attention-FFN Disaggregation for Modern MoE Models and Hardware Systems
por: Liu, Guowei, et al.
Publicado: (2026)
por: Liu, Guowei, et al.
Publicado: (2026)
AdapMoE: Adaptive Sensitivity-based Expert Gating and Management for Efficient MoE Inference
por: Zhong, Shuzhang, et al.
Publicado: (2024)
por: Zhong, Shuzhang, et al.
Publicado: (2024)
SMoE: An Algorithm-System Co-Design for Pushing MoE to the Edge via Expert Substitution
por: Zhu, Guoying, et al.
Publicado: (2025)
por: Zhu, Guoying, et al.
Publicado: (2025)
LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation
por: Shu, Fangxun, et al.
Publicado: (2024)
por: Shu, Fangxun, et al.
Publicado: (2024)
MoE-SpeQ: Speculative Quantized Decoding with Proactive Expert Prefetching and Offloading for Mixture-of-Experts
por: Wang, Wenfeng, et al.
Publicado: (2025)
por: Wang, Wenfeng, et al.
Publicado: (2025)
Sparsity is Combinatorial Depth: Quantifying MoE Expressivity via Tropical Geometry
por: Su, Ye, et al.
Publicado: (2026)
por: Su, Ye, et al.
Publicado: (2026)
Octavius: Mitigating Task Interference in MLLMs via LoRA-MoE
por: Chen, Zeren, et al.
Publicado: (2023)
por: Chen, Zeren, et al.
Publicado: (2023)
Routing Matters in MoE: Scaling Diffusion Transformers with Explicit Routing Guidance
por: Wei, Yujie, et al.
Publicado: (2025)
por: Wei, Yujie, et al.
Publicado: (2025)
AIMER: Calibration-Free Task-Agnostic MoE Pruning
por: Liu, Zongfang, et al.
Publicado: (2026)
por: Liu, Zongfang, et al.
Publicado: (2026)
TEAM: Temporal-Spatial Consistency Guided Expert Activation for MoE Diffusion Language Model Acceleration
por: Wei, Linye, et al.
Publicado: (2026)
por: Wei, Linye, et al.
Publicado: (2026)
Remoe: Towards Efficient and Low-Cost MoE Inference in Serverless Computing
por: Liu, Wentao, et al.
Publicado: (2025)
por: Liu, Wentao, et al.
Publicado: (2025)
UniMoE-Audio: Unified Speech and Music Generation with Dynamic-Capacity MoE
por: Liu, Zhenyu, et al.
Publicado: (2025)
por: Liu, Zhenyu, et al.
Publicado: (2025)
GRACE-MoE: Grouping and Replication with Locality-Aware Routing for Efficient Distributed MoE Inference
por: Han, Yu, et al.
Publicado: (2025)
por: Han, Yu, et al.
Publicado: (2025)
Does a Global Perspective Help Prune Sparse MoEs Elegantly?
por: Zhang, Zeliang, et al.
Publicado: (2026)
por: Zhang, Zeliang, et al.
Publicado: (2026)
Compositional-Degradation UAV Image Restoration: Conditional Decoupled MoE Network and A Benchmark
por: Yan, Jinquan, et al.
Publicado: (2026)
por: Yan, Jinquan, et al.
Publicado: (2026)
Condense, Don't Just Prune: Enhancing Efficiency and Performance in MoE Layer Pruning
por: Cao, Mingyu, et al.
Publicado: (2024)
por: Cao, Mingyu, et al.
Publicado: (2024)
Two Experts Are All You Need for Steering Thinking: Reinforcing Cognitive Effort in MoE Reasoning Models Without Additional Training
por: Wang, Mengru, et al.
Publicado: (2025)
por: Wang, Mengru, et al.
Publicado: (2025)
TAG-MoE: Task-Aware Gating for Unified Generative Mixture-of-Experts
por: Xu, Yu, et al.
Publicado: (2026)
por: Xu, Yu, et al.
Publicado: (2026)
Delta Decompression for MoE-based LLMs Compression
por: Gu, Hao, et al.
Publicado: (2025)
por: Gu, Hao, et al.
Publicado: (2025)
HOBBIT: A Mixed Precision Expert Offloading System for Fast MoE Inference
por: Tang, Peng, et al.
Publicado: (2024)
por: Tang, Peng, et al.
Publicado: (2024)
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models
por: Wang, Wei, et al.
Publicado: (2024)
por: Wang, Wei, et al.
Publicado: (2024)
Ejemplares similares
-
Pangu Pro MoE: Mixture of Grouped Experts for Efficient Sparsity
por: Tang, Yehui, et al.
Publicado: (2025) -
Pangu Ultra: Pushing the Limits of Dense Large Language Models on Ascend NPUs
por: Yin, Yichun, et al.
Publicado: (2025) -
MoE-Prism: Disentangling Monolithic Experts for Elastic MoE Services via Model-System Co-Designs
por: Xia, Xinfeng, et al.
Publicado: (2025) -
Alloc-MoE: Budget-Aware Expert Activation Allocation for Efficient Mixture-of-Experts Inference
por: Liu, Baihui, et al.
Publicado: (2026) -
LLaDA-MoE: A Sparse MoE Diffusion Language Model
por: Zhu, Fengqi, et al.
Publicado: (2025)