Marco-MoE: Open Multilingual Mixture-of-Expert Language Models with Efficient Upcycling
Fuente:
arXiv
Saved in:
| Main Authors: | Jiang, Fan, Zhao, Yu, Lyu, Chenyang, Shi, Tianqi, Du, Yichao, Jiang, Feihu, Wang, Longyue, Luo, Weihua |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Marco-o1: Towards Open Reasoning Models for Open-Ended Solutions
by: Zhao, Yu, et al.
Published: (2024)
by: Zhao, Yu, et al.
Published: (2024)
Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts
by: Li, Yunxin, et al.
Published: (2024)
by: Li, Yunxin, et al.
Published: (2024)
MoE-LPR: Multilingual Extension of Large Language Models through Mixture-of-Experts with Language Priors Routing
by: Zhou, Hao, et al.
Published: (2024)
by: Zhou, Hao, et al.
Published: (2024)
CLIP-MoE: Towards Building Mixture of Experts for CLIP with Diversified Multiplet Upcycling
by: Zhang, Jihai, et al.
Published: (2024)
by: Zhang, Jihai, et al.
Published: (2024)
MoE-Loco: Mixture of Experts for Multitask Locomotion
by: Huang, Runhan, et al.
Published: (2025)
by: Huang, Runhan, et al.
Published: (2025)
Marco-Bench-MIF: On Multilingual Instruction-Following Capability of Large Language Models
by: Zeng, Bo, et al.
Published: (2025)
by: Zeng, Bo, et al.
Published: (2025)
BLR-MoE: Boosted Language-Routing Mixture of Experts for Domain-Robust Multilingual E2E ASR
by: Ma, Guodong, et al.
Published: (2025)
by: Ma, Guodong, et al.
Published: (2025)
Hexa-MoE: Efficient and Heterogeneous-aware Training for Mixture-of-Experts
by: Luo, Shuqing, et al.
Published: (2024)
by: Luo, Shuqing, et al.
Published: (2024)
Marco DeepResearch: Unlocking Efficient Deep Research Agents via Verification-Centric Design
by: Zhu, Bin, et al.
Published: (2026)
by: Zhu, Bin, et al.
Published: (2026)
Marco-ASR: A Principled and Metric-Driven Framework for Fine-Tuning Large-Scale ASR Models for Domain Adaptation
by: Ni, Xuanfan, et al.
Published: (2025)
by: Ni, Xuanfan, et al.
Published: (2025)
UMEM: Unified Memory Extraction and Management Framework for Generalizable Memory
by: Ye, Yongshi, et al.
Published: (2026)
by: Ye, Yongshi, et al.
Published: (2026)
MoE-FFD: Mixture of Experts for Generalized and Parameter-Efficient Face Forgery Detection
by: Kong, Chenqi, et al.
Published: (2024)
by: Kong, Chenqi, et al.
Published: (2024)
Llama 3 Meets MoE: Efficient Upcycling
by: Vavre, Aditya, et al.
Published: (2024)
by: Vavre, Aditya, et al.
Published: (2024)
Horseshoe Mixtures-of-Experts (HS-MoE)
by: Polson, Nick, et al.
Published: (2026)
by: Polson, Nick, et al.
Published: (2026)
Pangu Pro MoE: Mixture of Grouped Experts for Efficient Sparsity
by: Tang, Yehui, et al.
Published: (2025)
by: Tang, Yehui, et al.
Published: (2025)
Marco-LLM: Bridging Languages via Massive Multilingual Training for Cross-Lingual Enhancement
by: Ming, Lingfeng, et al.
Published: (2024)
by: Ming, Lingfeng, et al.
Published: (2024)
Dirichlet-Prior Shaping: Guiding Expert Specialization in Upcycled MoEs
by: Mirvakhabova, Leyla, et al.
Published: (2025)
by: Mirvakhabova, Leyla, et al.
Published: (2025)
Med-MoE: Mixture of Domain-Specific Experts for Lightweight Medical Vision-Language Models
by: Jiang, Songtao, et al.
Published: (2024)
by: Jiang, Songtao, et al.
Published: (2024)
Adapted-MoE: Mixture of Experts with Test-Time Adaption for Anomaly Detection
by: Lei, Tianwu, et al.
Published: (2024)
by: Lei, Tianwu, et al.
Published: (2024)
OmniMoE: An Efficient MoE by Orchestrating Atomic Experts at Scale
by: Shi, Jingze, et al.
Published: (2026)
by: Shi, Jingze, et al.
Published: (2026)
DAG-MoE: From Simple Mixture to Structural Aggregation in Mixture-of-Experts
by: Feng, Jiarui, et al.
Published: (2026)
by: Feng, Jiarui, et al.
Published: (2026)
$\texttt{MoE-RBench}$: Towards Building Reliable Language Models with Sparse Mixture-of-Experts
by: Chen, Guanjie, et al.
Published: (2024)
by: Chen, Guanjie, et al.
Published: (2024)
MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
by: Lin, Bin, et al.
Published: (2024)
by: Lin, Bin, et al.
Published: (2024)
$\infty$-MoE: Generalizing Mixture of Experts to Infinite Experts
by: Takashiro, Shota, et al.
Published: (2026)
by: Takashiro, Shota, et al.
Published: (2026)
Grove MoE: Towards Efficient and Superior MoE LLMs with Adjugate Experts
by: Wu, Haoyuan, et al.
Published: (2025)
by: Wu, Haoyuan, et al.
Published: (2025)
Mixture of Experts (MoE): A Big Data Perspective
by: Gan, Wensheng, et al.
Published: (2025)
by: Gan, Wensheng, et al.
Published: (2025)
MoE-CAP: Benchmarking Cost, Accuracy and Performance of Sparse Mixture-of-Experts Systems
by: Jiang, Yinsicheng, et al.
Published: (2025)
by: Jiang, Yinsicheng, et al.
Published: (2025)
MoE-CAP: Benchmarking Cost, Accuracy and Performance of Sparse Mixture-of-Experts Systems
by: Jiang, Yinsicheng, et al.
Published: (2024)
by: Jiang, Yinsicheng, et al.
Published: (2024)
EPS-MoE: Expert Pipeline Scheduler for Cost-Efficient MoE Inference
by: Qian, Yulei, et al.
Published: (2024)
by: Qian, Yulei, et al.
Published: (2024)
Marco-o1 v2: Towards Widening The Distillation Bottleneck for Reasoning Models
by: Yin, Huifeng, et al.
Published: (2025)
by: Yin, Huifeng, et al.
Published: (2025)
MoIN: Mixture of Introvert Experts to Upcycle an LLM
by: Tejankar, Ajinkya, et al.
Published: (2024)
by: Tejankar, Ajinkya, et al.
Published: (2024)
MH-MoE: Multi-Head Mixture-of-Experts
by: Huang, Shaohan, et al.
Published: (2024)
by: Huang, Shaohan, et al.
Published: (2024)
CuMo: Scaling Multimodal LLM with Co-Upcycled Mixture-of-Experts
by: Li, Jiachen, et al.
Published: (2024)
by: Li, Jiachen, et al.
Published: (2024)
Steer-MoE: Efficient Audio-Language Alignment with a Mixture-of-Experts Steering Module
by: Feng, Ruitao, et al.
Published: (2025)
by: Feng, Ruitao, et al.
Published: (2025)
SiDA-MoE: Sparsity-Inspired Data-Aware Serving for Efficient and Scalable Large Mixture-of-Experts Models
by: Du, Zhixu, et al.
Published: (2023)
by: Du, Zhixu, et al.
Published: (2023)
MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache
by: Xue, Leyang, et al.
Published: (2024)
by: Xue, Leyang, et al.
Published: (2024)
Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient
by: Ludziejewski, Jan, et al.
Published: (2025)
by: Ludziejewski, Jan, et al.
Published: (2025)
MoE-Mamba: Efficient Selective State Space Models with Mixture of Experts
by: Pióro, Maciej, et al.
Published: (2024)
by: Pióro, Maciej, et al.
Published: (2024)
MoE++: Accelerating Mixture-of-Experts Methods with Zero-Computation Experts
by: Jin, Peng, et al.
Published: (2024)
by: Jin, Peng, et al.
Published: (2024)
RS-MoE: A Vision-Language Model with Mixture of Experts for Remote Sensing Image Captioning and Visual Question Answering
by: Lin, Hui, et al.
Published: (2024)
by: Lin, Hui, et al.
Published: (2024)
Similar Items
-
Marco-o1: Towards Open Reasoning Models for Open-Ended Solutions
by: Zhao, Yu, et al.
Published: (2024) -
Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts
by: Li, Yunxin, et al.
Published: (2024) -
MoE-LPR: Multilingual Extension of Large Language Models through Mixture-of-Experts with Language Priors Routing
by: Zhou, Hao, et al.
Published: (2024) -
CLIP-MoE: Towards Building Mixture of Experts for CLIP with Diversified Multiplet Upcycling
by: Zhang, Jihai, et al.
Published: (2024) -
MoE-Loco: Mixture of Experts for Multitask Locomotion
by: Huang, Runhan, et al.
Published: (2025)