DELLA-Merging: Reducing Interference in Model Merging through Magnitude-Based Sampling

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Deep, Pala Tej, Bhardwaj, Rishabh, Poria, Soujanya
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909225477734400
author Deep, Pala Tej
Bhardwaj, Rishabh
Poria, Soujanya
author_facet Deep, Pala Tej
Bhardwaj, Rishabh
Poria, Soujanya
contents With the proliferation of domain-specific models, model merging has emerged as a set of techniques that combine the capabilities of multiple models into one that can multitask without the cost of additional training. In this paper, we propose a new model merging technique, Drop and rEscaLe via sampLing with mAgnitude (DELLA-Merging), that employs a novel pruning technique, MAGPRUNE, which shows significant advantages over DARE and TIES. MAGPRUNE first ranks the parameters in order of their magnitude and assigns higher dropout probabilities (p) to parameters with lower ranks corresponding to lower magnitudes. To approximate the original embeddings, MAGPRUNE employs a rescaling operation on the parameters that survive the random dropping by 1/(1 - p). On three different expert models considered for merging (LM, Math, Code) and corresponding benchmark datasets (AlpacaEval, GSM8K, MBPP), DELLA shows an average improvement of 2.4 points over baseline methods employing delta parameter pruning (an improvement of 3.6 points over TIES, 1.2 points over DARE), and 11.1 points over the no-pruning baseline (TA). We release the source code at: https://github.com/declare-lab/della.
format Preprint
id arxiv_https___arxiv_org_abs_2406_11617
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DELLA-Merging: Reducing Interference in Model Merging through Magnitude-Based Sampling
Deep, Pala Tej
Bhardwaj, Rishabh
Poria, Soujanya
Computation and Language
With the proliferation of domain-specific models, model merging has emerged as a set of techniques that combine the capabilities of multiple models into one that can multitask without the cost of additional training. In this paper, we propose a new model merging technique, Drop and rEscaLe via sampLing with mAgnitude (DELLA-Merging), that employs a novel pruning technique, MAGPRUNE, which shows significant advantages over DARE and TIES. MAGPRUNE first ranks the parameters in order of their magnitude and assigns higher dropout probabilities (p) to parameters with lower ranks corresponding to lower magnitudes. To approximate the original embeddings, MAGPRUNE employs a rescaling operation on the parameters that survive the random dropping by 1/(1 - p). On three different expert models considered for merging (LM, Math, Code) and corresponding benchmark datasets (AlpacaEval, GSM8K, MBPP), DELLA shows an average improvement of 2.4 points over baseline methods employing delta parameter pruning (an improvement of 3.6 points over TIES, 1.2 points over DARE), and 11.1 points over the no-pruning baseline (TA). We release the source code at: https://github.com/declare-lab/della.
title DELLA-Merging: Reducing Interference in Model Merging through Magnitude-Based Sampling
topic Computation and Language
url https://arxiv.org/abs/2406.11617