Revisiting Weight Averaging for Model Merging

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Choi, Jiho, Kim, Donggyun, Lee, Chanhyuk, Hong, Seunghoon
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910902270296064
author Choi, Jiho
Kim, Donggyun
Lee, Chanhyuk
Hong, Seunghoon
author_facet Choi, Jiho
Kim, Donggyun
Lee, Chanhyuk
Hong, Seunghoon
contents Model merging aims to build a multi-task learner by combining the parameters of individually fine-tuned models without additional training. While a straightforward approach is to average model parameters across tasks, this often results in suboptimal performance due to interference among parameters across tasks. In this paper, we present intriguing results that weight averaging implicitly induces task vectors centered around the weight averaging itself and that applying a low-rank approximation to these centered task vectors significantly improves merging performance. Our analysis shows that centering the task vectors effectively reduces task interference and most of task-specific knowledge is concentrated in the top singular vectors. Our method demonstrates robust and scalable performance on vision benchmarks across varying numbers of tasks and model sizes. Furthermore, we observe that our approach is applicable to natural language processing tasks with competitive performance.
format Preprint
id arxiv_https___arxiv_org_abs_2412_12153
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Revisiting Weight Averaging for Model Merging
Choi, Jiho
Kim, Donggyun
Lee, Chanhyuk
Hong, Seunghoon
Machine Learning
Artificial Intelligence
Model merging aims to build a multi-task learner by combining the parameters of individually fine-tuned models without additional training. While a straightforward approach is to average model parameters across tasks, this often results in suboptimal performance due to interference among parameters across tasks. In this paper, we present intriguing results that weight averaging implicitly induces task vectors centered around the weight averaging itself and that applying a low-rank approximation to these centered task vectors significantly improves merging performance. Our analysis shows that centering the task vectors effectively reduces task interference and most of task-specific knowledge is concentrated in the top singular vectors. Our method demonstrates robust and scalable performance on vision benchmarks across varying numbers of tasks and model sizes. Furthermore, we observe that our approach is applicable to natural language processing tasks with competitive performance.
title Revisiting Weight Averaging for Model Merging
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2412.12153