Skywork-MoE: A Deep Dive into Training Techniques for Mixture-of-Experts Language Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Wei, Tianwen, Zhu, Bo, Zhao, Liang, Cheng, Cheng, Li, Biye, Lü, Weiwei, Cheng, Peng, Zhang, Jianhao, Zhang, Xiaoyu, Zeng, Liang, Wang, Xiaokun, Ma, Yutuan, Hu, Rui, Yan, Shuicheng, Fang, Han, Zhou, Yahui |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
LongSkywork: A Training Recipe for Efficiently Extending Context Length in Large Language Models
par: Zhao, Liang, et autres
Publié: (2024)
par: Zhao, Liang, et autres
Publié: (2024)
Skywork-Math: Data Scaling Laws for Mathematical Reasoning in Large Language Models -- The Story Goes On
par: Zeng, Liang, et autres
Publié: (2024)
par: Zeng, Liang, et autres
Publié: (2024)
Skywork Open Reasoner 1 Technical Report
par: He, Jujie, et autres
Publié: (2025)
par: He, Jujie, et autres
Publié: (2025)
MoE++: Accelerating Mixture-of-Experts Methods with Zero-Computation Experts
par: Jin, Peng, et autres
Publié: (2024)
par: Jin, Peng, et autres
Publié: (2024)
MoE-DisCo:Low Economy Cost Training Mixture-of-Experts Models
par: Ye, Xin, et autres
Publié: (2026)
par: Ye, Xin, et autres
Publié: (2026)
Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs
par: Liu, Chris Yuhao, et autres
Publié: (2024)
par: Liu, Chris Yuhao, et autres
Publié: (2024)
CLIP-MoE: Towards Building Mixture of Experts for CLIP with Diversified Multiplet Upcycling
par: Zhang, Jihai, et autres
Publié: (2024)
par: Zhang, Jihai, et autres
Publié: (2024)
Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs
par: Zeng, Liang, et autres
Publié: (2025)
par: Zeng, Liang, et autres
Publié: (2025)
Linear-MoE: Linear Sequence Modeling Meets Mixture-of-Experts
par: Sun, Weigao, et autres
Publié: (2025)
par: Sun, Weigao, et autres
Publié: (2025)
EAC-MoE: Expert-Selection Aware Compressor for Mixture-of-Experts Large Language Models
par: Chen, Yuanteng, et autres
Publié: (2025)
par: Chen, Yuanteng, et autres
Publié: (2025)
Point-MoE: Large-Scale Multi-Dataset Training with Mixture-of-Experts for 3D Semantic Segmentation
par: Chen, Xuweiyi, et autres
Publié: (2025)
par: Chen, Xuweiyi, et autres
Publié: (2025)
MoE-I$^2$: Compressing Mixture of Experts Models through Inter-Expert Pruning and Intra-Expert Low-Rank Decomposition
par: Yang, Cheng, et autres
Publié: (2024)
par: Yang, Cheng, et autres
Publié: (2024)
MoPE: Mixture of Prefix Experts for Zero-Shot Dialogue State Tracking
par: Tang, Tianwen, et autres
Publié: (2024)
par: Tang, Tianwen, et autres
Publié: (2024)
Steer-MoE: Efficient Audio-Language Alignment with a Mixture-of-Experts Steering Module
par: Feng, Ruitao, et autres
Publié: (2025)
par: Feng, Ruitao, et autres
Publié: (2025)
$\texttt{MoE-RBench}$: Towards Building Reliable Language Models with Sparse Mixture-of-Experts
par: Chen, Guanjie, et autres
Publié: (2024)
par: Chen, Guanjie, et autres
Publié: (2024)
DAG-MoE: From Simple Mixture to Structural Aggregation in Mixture-of-Experts
par: Feng, Jiarui, et autres
Publié: (2026)
par: Feng, Jiarui, et autres
Publié: (2026)
TAG-MoE: Task-Aware Gating for Unified Generative Mixture-of-Experts
par: Xu, Yu, et autres
Publié: (2026)
par: Xu, Yu, et autres
Publié: (2026)
Skywork R1V2: Multimodal Hybrid Reinforcement Learning for Reasoning
par: Wang, Peiyu, et autres
Publié: (2025)
par: Wang, Peiyu, et autres
Publié: (2025)
MoBE: Mixture-of-Basis-Experts for Compressing MoE-based LLMs
par: Chen, Xiaodong, et autres
Publié: (2025)
par: Chen, Xiaodong, et autres
Publié: (2025)
Hexa-MoE: Efficient and Heterogeneous-aware Training for Mixture-of-Experts
par: Luo, Shuqing, et autres
Publié: (2024)
par: Luo, Shuqing, et autres
Publié: (2024)
LLaMA-MoE v2: Exploring Sparsity of LLaMA from Perspective of Mixture-of-Experts with Post-Training
par: Qu, Xiaoye, et autres
Publié: (2024)
par: Qu, Xiaoye, et autres
Publié: (2024)
Horseshoe Mixtures-of-Experts (HS-MoE)
par: Polson, Nick, et autres
Publié: (2026)
par: Polson, Nick, et autres
Publié: (2026)
AW-MoE: All-Weather Mixture of Experts for Robust Multi-Modal 3D Object Detection
par: Lin, Hongwei, et autres
Publié: (2026)
par: Lin, Hongwei, et autres
Publié: (2026)
EPS-MoE: Expert Pipeline Scheduler for Cost-Efficient MoE Inference
par: Qian, Yulei, et autres
Publié: (2024)
par: Qian, Yulei, et autres
Publié: (2024)
Skywork-R1V3 Technical Report
par: Shen, Wei, et autres
Publié: (2025)
par: Shen, Wei, et autres
Publié: (2025)
PC-MoE: Memory-Efficient and Privacy-Preserving Collaborative Training for Mixture-of-Experts LLMs
par: Zhang, Ze Yu, et autres
Publié: (2025)
par: Zhang, Ze Yu, et autres
Publié: (2025)
FLEX-MoE: Federated Mixture-of-Experts with Load-balanced Expert Assignment for Edge Computing
par: Zhang, Boyang, et autres
Publié: (2025)
par: Zhang, Boyang, et autres
Publié: (2025)
$\infty$-MoE: Generalizing Mixture of Experts to Infinite Experts
par: Takashiro, Shota, et autres
Publié: (2026)
par: Takashiro, Shota, et autres
Publié: (2026)
X-MoE: Enabling Scalable Training for Emerging Mixture-of-Experts Architectures on HPC Platforms
par: Yuan, Yueming, et autres
Publié: (2025)
par: Yuan, Yueming, et autres
Publié: (2025)
LAER-MoE: Load-Adaptive Expert Re-layout for Efficient Mixture-of-Experts Training
par: Liu, Xinyi, et autres
Publié: (2026)
par: Liu, Xinyi, et autres
Publié: (2026)
Skywork UniPic: Unified Autoregressive Modeling for Visual Understanding and Generation
par: Wang, Peiyu, et autres
Publié: (2025)
par: Wang, Peiyu, et autres
Publié: (2025)
MoE Jetpack: From Dense Checkpoints to Adaptive Mixture of Experts for Vision Tasks
par: Zhu, Xingkui, et autres
Publié: (2024)
par: Zhu, Xingkui, et autres
Publié: (2024)
LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-training
par: Zhu, Tong, et autres
Publié: (2024)
par: Zhu, Tong, et autres
Publié: (2024)
MH-MoE: Multi-Head Mixture-of-Experts
par: Huang, Shaohan, et autres
Publié: (2024)
par: Huang, Shaohan, et autres
Publié: (2024)
MoE-Loco: Mixture of Experts for Multitask Locomotion
par: Huang, Runhan, et autres
Publié: (2025)
par: Huang, Runhan, et autres
Publié: (2025)
MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts
par: Xue, Heyang, et autres
Publié: (2025)
par: Xue, Heyang, et autres
Publié: (2025)
Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts
par: Li, Yunxin, et autres
Publié: (2024)
par: Li, Yunxin, et autres
Publié: (2024)
GM-MoE: Low-Light Enhancement with Gated-Mechanism Mixture-of-Experts
par: Liao, Minwen, et autres
Publié: (2025)
par: Liao, Minwen, et autres
Publié: (2025)
Elastic MoE: Unlocking the Inference-Time Scalability of Mixture-of-Experts
par: Gu, Naibin, et autres
Publié: (2025)
par: Gu, Naibin, et autres
Publié: (2025)
Moirai-MoE: Empowering Time Series Foundation Models with Sparse Mixture of Experts
par: Liu, Xu, et autres
Publié: (2024)
par: Liu, Xu, et autres
Publié: (2024)
Documents similaires
-
LongSkywork: A Training Recipe for Efficiently Extending Context Length in Large Language Models
par: Zhao, Liang, et autres
Publié: (2024) -
Skywork-Math: Data Scaling Laws for Mathematical Reasoning in Large Language Models -- The Story Goes On
par: Zeng, Liang, et autres
Publié: (2024) -
Skywork Open Reasoner 1 Technical Report
par: He, Jujie, et autres
Publié: (2025) -
MoE++: Accelerating Mixture-of-Experts Methods with Zero-Computation Experts
par: Jin, Peng, et autres
Publié: (2024) -
MoE-DisCo:Low Economy Cost Training Mixture-of-Experts Models
par: Ye, Xin, et autres
Publié: (2026)