Exploring Sparse Adapters for Scalable Merging of Parameter Efficient Experts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Arnob, Samin Yeasar, Su, Zhan, Kim, Minseon, Ostapenko, Oleksiy, Ohib, Riyasat, Saleh, Esra'a, Precup, Doina, Caccia, Lucas, Sordoni, Alessandro
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911053678379008
author Arnob, Samin Yeasar
Su, Zhan
Kim, Minseon
Ostapenko, Oleksiy
Ohib, Riyasat
Saleh, Esra'a
Precup, Doina
Caccia, Lucas
Sordoni, Alessandro
author_facet Arnob, Samin Yeasar
Su, Zhan
Kim, Minseon
Ostapenko, Oleksiy
Ohib, Riyasat
Saleh, Esra'a
Precup, Doina
Caccia, Lucas
Sordoni, Alessandro
contents Merging parameter-efficient task experts has recently gained growing attention as a way to build modular architectures that can be rapidly adapted on the fly for specific downstream tasks, without requiring additional fine-tuning. Typically, LoRA serves as the foundational building block of such parameter-efficient modular architectures, leveraging low-rank weight structures to reduce the number of trainable parameters. In this paper, we study the properties of sparse adapters, which train only a subset of weights in the base neural network, as potential building blocks of modular architectures. First, we propose a simple method for training highly effective sparse adapters, which is conceptually simpler than existing methods in the literature and surprisingly outperforms both LoRA and full fine-tuning in our setting. Next, we investigate the merging properties of these sparse adapters by merging adapters for up to 20 natural language processing tasks, thus scaling beyond what is usually studied in the literature. Our findings demonstrate that sparse adapters yield superior in-distribution performance post-merging compared to LoRA or full model merging. Achieving strong held-out performance remains a challenge for all methods considered.
format Preprint
id arxiv_https___arxiv_org_abs_2507_07140
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Exploring Sparse Adapters for Scalable Merging of Parameter Efficient Experts
Arnob, Samin Yeasar
Su, Zhan
Kim, Minseon
Ostapenko, Oleksiy
Ohib, Riyasat
Saleh, Esra'a
Precup, Doina
Caccia, Lucas
Sordoni, Alessandro
Machine Learning
Merging parameter-efficient task experts has recently gained growing attention as a way to build modular architectures that can be rapidly adapted on the fly for specific downstream tasks, without requiring additional fine-tuning. Typically, LoRA serves as the foundational building block of such parameter-efficient modular architectures, leveraging low-rank weight structures to reduce the number of trainable parameters. In this paper, we study the properties of sparse adapters, which train only a subset of weights in the base neural network, as potential building blocks of modular architectures. First, we propose a simple method for training highly effective sparse adapters, which is conceptually simpler than existing methods in the literature and surprisingly outperforms both LoRA and full fine-tuning in our setting. Next, we investigate the merging properties of these sparse adapters by merging adapters for up to 20 natural language processing tasks, thus scaling beyond what is usually studied in the literature. Our findings demonstrate that sparse adapters yield superior in-distribution performance post-merging compared to LoRA or full model merging. Achieving strong held-out performance remains a challenge for all methods considered.
title Exploring Sparse Adapters for Scalable Merging of Parameter Efficient Experts
topic Machine Learning
url https://arxiv.org/abs/2507.07140