LD-MoLE: Learnable Dynamic Routing for Mixture of LoRA Experts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhuang, Yuan, Shen, Yi, Bian, Yuexin, Su, Qing, Ji, Shihao, Shi, Yuanyuan, Miao, Fei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912922154827776
author Zhuang, Yuan
Shen, Yi
Bian, Yuexin
Su, Qing
Ji, Shihao
Shi, Yuanyuan
Miao, Fei
author_facet Zhuang, Yuan
Shen, Yi
Bian, Yuexin
Su, Qing
Ji, Shihao
Shi, Yuanyuan
Miao, Fei
contents Recent studies have shown that combining parameter-efficient fine-tuning (PEFT) with mixture-of-experts (MoE) is an effective strategy for adapting large language models (LLMs) to the downstream tasks. However, most existing approaches rely on conventional TopK routing, which requires careful hyperparameter tuning and assigns a fixed number of experts to each token. In this work, we propose LD-MoLE, a Learnable Dynamic routing mechanism for Mixture of LoRA Experts that enables adaptive, token-dependent, and layer-wise expert allocation. Our method replaces the non-differentiable TopK selection with a differentiable routing function and a closed-form solution. Moreover, our design allows the model to adaptively determine the number of experts to activate for each token at different layers. In addition, we introduce an analytical sparsity control objective to regularize the number of activated experts. Extensive experiments on the Qwen3-1.7B and Llama-3.2-3B models show that LD-MoLE achieves the highest average scores compared to state-of-the-art baselines, across a diverse set of benchmarks. Our method not only achieves superior performance, but also demonstrates the ability to learn token-dependent and layer-wise expert allocation.
format Preprint
id arxiv_https___arxiv_org_abs_2509_25684
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LD-MoLE: Learnable Dynamic Routing for Mixture of LoRA Experts
Zhuang, Yuan
Shen, Yi
Bian, Yuexin
Su, Qing
Ji, Shihao
Shi, Yuanyuan
Miao, Fei
Computation and Language
Artificial Intelligence
Recent studies have shown that combining parameter-efficient fine-tuning (PEFT) with mixture-of-experts (MoE) is an effective strategy for adapting large language models (LLMs) to the downstream tasks. However, most existing approaches rely on conventional TopK routing, which requires careful hyperparameter tuning and assigns a fixed number of experts to each token. In this work, we propose LD-MoLE, a Learnable Dynamic routing mechanism for Mixture of LoRA Experts that enables adaptive, token-dependent, and layer-wise expert allocation. Our method replaces the non-differentiable TopK selection with a differentiable routing function and a closed-form solution. Moreover, our design allows the model to adaptively determine the number of experts to activate for each token at different layers. In addition, we introduce an analytical sparsity control objective to regularize the number of activated experts. Extensive experiments on the Qwen3-1.7B and Llama-3.2-3B models show that LD-MoLE achieves the highest average scores compared to state-of-the-art baselines, across a diverse set of benchmarks. Our method not only achieves superior performance, but also demonstrates the ability to learn token-dependent and layer-wise expert allocation.
title LD-MoLE: Learnable Dynamic Routing for Mixture of LoRA Experts
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2509.25684