Router Upcycling: Leveraging Mixture-of-Routers in Mixture-of-Experts Upcycling
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866915472389177344 |
|---|---|
| author | Ran, Junfeng Zhao, Guangxiang Wu, Yuhan Zhu, Dawei Wu, Longyun Zhao, Yikai Yang, Tong Sun, Lin Zhang, Xiangzheng Li, Sujian |
| author_facet | Ran, Junfeng Zhao, Guangxiang Wu, Yuhan Zhu, Dawei Wu, Longyun Zhao, Yikai Yang, Tong Sun, Lin Zhang, Xiangzheng Li, Sujian |
| contents | The Mixture-of-Experts (MoE) models have gained significant attention in deep learning due to their dynamic resource allocation and superior performance across diverse tasks. However, efficiently training these models remains challenging. The MoE upcycling technique has been proposed to reuse and improve existing model components, thereby minimizing training overhead. Despite this, simple routers, such as linear routers, often struggle with complex routing tasks within MoE upcycling. In response, we propose a novel routing technique called Router Upcycling to enhance the performance of MoE upcycling models. Our approach initializes multiple routers from the attention heads of preceding attention layers during upcycling. These routers collaboratively assign tokens to specialized experts in an attention-like manner. Each token is processed into diverse queries and aligned with the experts' features (serving as keys). Experimental results demonstrate that our method achieves state-of-the-art (SOTA) performance, outperforming other upcycling baselines. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_00679 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Router Upcycling: Leveraging Mixture-of-Routers in Mixture-of-Experts Upcycling Ran, Junfeng Zhao, Guangxiang Wu, Yuhan Zhu, Dawei Wu, Longyun Zhao, Yikai Yang, Tong Sun, Lin Zhang, Xiangzheng Li, Sujian Computation and Language The Mixture-of-Experts (MoE) models have gained significant attention in deep learning due to their dynamic resource allocation and superior performance across diverse tasks. However, efficiently training these models remains challenging. The MoE upcycling technique has been proposed to reuse and improve existing model components, thereby minimizing training overhead. Despite this, simple routers, such as linear routers, often struggle with complex routing tasks within MoE upcycling. In response, we propose a novel routing technique called Router Upcycling to enhance the performance of MoE upcycling models. Our approach initializes multiple routers from the attention heads of preceding attention layers during upcycling. These routers collaboratively assign tokens to specialized experts in an attention-like manner. Each token is processed into diverse queries and aligned with the experts' features (serving as keys). Experimental results demonstrate that our method achieves state-of-the-art (SOTA) performance, outperforming other upcycling baselines. |
| title | Router Upcycling: Leveraging Mixture-of-Routers in Mixture-of-Experts Upcycling |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2509.00679 |