S'MoRE: Structural Mixture of Residual Experts for Parameter-Efficient LLM Fine-tuning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zeng, Hanqing, Xia, Yinglong, Zhao, Zhuokai, Jiang, Chuan, Zhang, Qiang, Liu, Jiayi, Zhang, Qunshu, Zhang, Lizhu, Fan, Xiangjun, Zhang, Benyu
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917048621203456
author Zeng, Hanqing
Xia, Yinglong
Zhao, Zhuokai
Jiang, Chuan
Zhang, Qiang
Liu, Jiayi
Zhang, Qunshu
Zhang, Lizhu
Fan, Xiangjun
Zhang, Benyu
author_facet Zeng, Hanqing
Xia, Yinglong
Zhao, Zhuokai
Jiang, Chuan
Zhang, Qiang
Liu, Jiayi
Zhang, Qunshu
Zhang, Lizhu
Fan, Xiangjun
Zhang, Benyu
contents Fine-tuning pre-trained large language models (LLMs) presents a dual challenge of balancing parameter efficiency and model capacity. Existing methods like low-rank adaptations (LoRA) are efficient but lack flexibility, while Mixture-of-Experts (MoE) enhance model capacity at the cost of more & under-utilized parameters. To address these limitations, we propose Structural Mixture of Residual Experts (S'MoRE), a novel framework that seamlessly integrates the efficiency of LoRA with the flexibility of MoE. Conceptually, S'MoRE employs hierarchical low-rank decomposition of expert weights, yielding residuals of varying orders interconnected in a multi-layer structure. By routing input tokens through sub-trees of residuals, S'MoRE emulates the capacity of numerous experts by instantiating and assembling just a few low-rank matrices. We craft the inter-layer propagation of S'MoRE's residuals as a special type of Graph Neural Network (GNN), and prove that under similar parameter budget, S'MoRE improves structural flexibility of traditional MoE (or Mixture-of-LoRA) by exponential order. Comprehensive theoretical analysis and empirical results demonstrate that S'MoRE achieves superior fine-tuning performance, offering a transformative approach for efficient LLM adaptation. Our implementation is available at: https://github.com/ZimpleX/SMoRE-LLM.
format Preprint
id arxiv_https___arxiv_org_abs_2504_06426
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle S'MoRE: Structural Mixture of Residual Experts for Parameter-Efficient LLM Fine-tuning
Zeng, Hanqing
Xia, Yinglong
Zhao, Zhuokai
Jiang, Chuan
Zhang, Qiang
Liu, Jiayi
Zhang, Qunshu
Zhang, Lizhu
Fan, Xiangjun
Zhang, Benyu
Computation and Language
Machine Learning
Fine-tuning pre-trained large language models (LLMs) presents a dual challenge of balancing parameter efficiency and model capacity. Existing methods like low-rank adaptations (LoRA) are efficient but lack flexibility, while Mixture-of-Experts (MoE) enhance model capacity at the cost of more & under-utilized parameters. To address these limitations, we propose Structural Mixture of Residual Experts (S'MoRE), a novel framework that seamlessly integrates the efficiency of LoRA with the flexibility of MoE. Conceptually, S'MoRE employs hierarchical low-rank decomposition of expert weights, yielding residuals of varying orders interconnected in a multi-layer structure. By routing input tokens through sub-trees of residuals, S'MoRE emulates the capacity of numerous experts by instantiating and assembling just a few low-rank matrices. We craft the inter-layer propagation of S'MoRE's residuals as a special type of Graph Neural Network (GNN), and prove that under similar parameter budget, S'MoRE improves structural flexibility of traditional MoE (or Mixture-of-LoRA) by exponential order. Comprehensive theoretical analysis and empirical results demonstrate that S'MoRE achieves superior fine-tuning performance, offering a transformative approach for efficient LLM adaptation. Our implementation is available at: https://github.com/ZimpleX/SMoRE-LLM.
title S'MoRE: Structural Mixture of Residual Experts for Parameter-Efficient LLM Fine-tuning
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2504.06426