Router Upcycling: Leveraging Mixture-of-Routers in Mixture-of-Experts Upcycling

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ran, Junfeng, Zhao, Guangxiang, Wu, Yuhan, Zhu, Dawei, Wu, Longyun, Zhao, Yikai, Yang, Tong, Sun, Lin, Zhang, Xiangzheng, Li, Sujian
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915472389177344
author Ran, Junfeng
Zhao, Guangxiang
Wu, Yuhan
Zhu, Dawei
Wu, Longyun
Zhao, Yikai
Yang, Tong
Sun, Lin
Zhang, Xiangzheng
Li, Sujian
author_facet Ran, Junfeng
Zhao, Guangxiang
Wu, Yuhan
Zhu, Dawei
Wu, Longyun
Zhao, Yikai
Yang, Tong
Sun, Lin
Zhang, Xiangzheng
Li, Sujian
contents The Mixture-of-Experts (MoE) models have gained significant attention in deep learning due to their dynamic resource allocation and superior performance across diverse tasks. However, efficiently training these models remains challenging. The MoE upcycling technique has been proposed to reuse and improve existing model components, thereby minimizing training overhead. Despite this, simple routers, such as linear routers, often struggle with complex routing tasks within MoE upcycling. In response, we propose a novel routing technique called Router Upcycling to enhance the performance of MoE upcycling models. Our approach initializes multiple routers from the attention heads of preceding attention layers during upcycling. These routers collaboratively assign tokens to specialized experts in an attention-like manner. Each token is processed into diverse queries and aligned with the experts' features (serving as keys). Experimental results demonstrate that our method achieves state-of-the-art (SOTA) performance, outperforming other upcycling baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2509_00679
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Router Upcycling: Leveraging Mixture-of-Routers in Mixture-of-Experts Upcycling
Ran, Junfeng
Zhao, Guangxiang
Wu, Yuhan
Zhu, Dawei
Wu, Longyun
Zhao, Yikai
Yang, Tong
Sun, Lin
Zhang, Xiangzheng
Li, Sujian
Computation and Language
The Mixture-of-Experts (MoE) models have gained significant attention in deep learning due to their dynamic resource allocation and superior performance across diverse tasks. However, efficiently training these models remains challenging. The MoE upcycling technique has been proposed to reuse and improve existing model components, thereby minimizing training overhead. Despite this, simple routers, such as linear routers, often struggle with complex routing tasks within MoE upcycling. In response, we propose a novel routing technique called Router Upcycling to enhance the performance of MoE upcycling models. Our approach initializes multiple routers from the attention heads of preceding attention layers during upcycling. These routers collaboratively assign tokens to specialized experts in an attention-like manner. Each token is processed into diverse queries and aligned with the experts' features (serving as keys). Experimental results demonstrate that our method achieves state-of-the-art (SOTA) performance, outperforming other upcycling baselines.
title Router Upcycling: Leveraging Mixture-of-Routers in Mixture-of-Experts Upcycling
topic Computation and Language
url https://arxiv.org/abs/2509.00679