Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Mu, Lin, Wang, Haiyang, Ni, Li, Sang, Lei, Wu, Zhize, Jin, Peiquan, Zhang, Yiwen
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2604.06291
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911586526953472
author Mu, Lin
Wang, Haiyang
Ni, Li
Sang, Lei
Wu, Zhize
Jin, Peiquan
Zhang, Yiwen
author_facet Mu, Lin
Wang, Haiyang
Ni, Li
Sang, Lei
Wu, Zhize
Jin, Peiquan
Zhang, Yiwen
contents Low-Rank Adaptation (LoRA) enables parameter-efficient fine-tuning of Large Language Models (LLMs), and recent Mixture-of-Experts (MoE) extensions further enhance flexibility by dynamically combining multiple LoRA experts. However, existing MoE-augmented LoRA methods assume that experts operate independently, often leading to unstable routing, expert dominance. In this paper, we propose \textbf{TalkLoRA}, a communication-aware MoELoRA framework that relaxes this independence assumption by introducing expert-level communication prior to routing. TalkLoRA equips low-rank experts with a lightweight Talking Module that enables controlled information exchange across expert subspaces, producing a more robust global signal for routing. Theoretically, we show that expert communication smooths routing dynamics by mitigating perturbation amplification while strictly generalizing existing MoELoRA architectures. Empirically, TalkLoRA consistently outperforms vanilla LoRA and MoELoRA across diverse language understanding and generation tasks, achieving higher parameter efficiency and more balanced expert routing under comparable parameter budgets. These results highlight structured expert communication as a principled and effective enhancement for MoE-based parameter-efficient adaptation. Code is available at https://github.com/why0129/TalkLoRA.
format Preprint
id arxiv_https___arxiv_org_abs_2604_06291
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle TalkLoRA: Communication-Aware Mixture of Low-Rank Adaptation for Large Language Models
Mu, Lin
Wang, Haiyang
Ni, Li
Sang, Lei
Wu, Zhize
Jin, Peiquan
Zhang, Yiwen
Machine Learning
Artificial Intelligence
Low-Rank Adaptation (LoRA) enables parameter-efficient fine-tuning of Large Language Models (LLMs), and recent Mixture-of-Experts (MoE) extensions further enhance flexibility by dynamically combining multiple LoRA experts. However, existing MoE-augmented LoRA methods assume that experts operate independently, often leading to unstable routing, expert dominance. In this paper, we propose \textbf{TalkLoRA}, a communication-aware MoELoRA framework that relaxes this independence assumption by introducing expert-level communication prior to routing. TalkLoRA equips low-rank experts with a lightweight Talking Module that enables controlled information exchange across expert subspaces, producing a more robust global signal for routing. Theoretically, we show that expert communication smooths routing dynamics by mitigating perturbation amplification while strictly generalizing existing MoELoRA architectures. Empirically, TalkLoRA consistently outperforms vanilla LoRA and MoELoRA across diverse language understanding and generation tasks, achieving higher parameter efficiency and more balanced expert routing under comparable parameter budgets. These results highlight structured expert communication as a principled and effective enhancement for MoE-based parameter-efficient adaptation. Code is available at https://github.com/why0129/TalkLoRA.
title TalkLoRA: Communication-Aware Mixture of Low-Rank Adaptation for Large Language Models
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2604.06291