Subspace-Boosted Model Merging

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Skorobogat, Ronald, Roth, Karsten, Georgescu, Mariana-Iuliana
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909972931018752
author Skorobogat, Ronald
Roth, Karsten
Georgescu, Mariana-Iuliana
author_facet Skorobogat, Ronald
Roth, Karsten
Georgescu, Mariana-Iuliana
contents Model merging enables the combination of multiple specialized expert models into a single model capable of performing multiple tasks. However, the benefits of merging an increasing amount of specialized experts generally lead to diminishing returns and reduced overall performance gains. In this work, we empirically and theoretically analyze this limitation, proving that for Task Arithmetic-based methods, as more experts are merged, the common information dominates the task-specific information, leading to inevitable rank collapse. To mitigate this issue, we introduce Subspace Boosting, which operates on the singular value decomposed task vector space and maintains task vector ranks. Subspace Boosting raises merging efficacy for up to 20 experts by large margins of more than 10% when evaluated on both vision and language benchmarks. Moreover, we propose employing Higher-Order Generalized Singular Value Decomposition to quantify task similarity, offering a new interpretable perspective on model merging. Code and models are available at https://github.com/ronskoro/Subspace-Boosting.
format Preprint
id arxiv_https___arxiv_org_abs_2506_16506
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Subspace-Boosted Model Merging
Skorobogat, Ronald
Roth, Karsten
Georgescu, Mariana-Iuliana
Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
Model merging enables the combination of multiple specialized expert models into a single model capable of performing multiple tasks. However, the benefits of merging an increasing amount of specialized experts generally lead to diminishing returns and reduced overall performance gains. In this work, we empirically and theoretically analyze this limitation, proving that for Task Arithmetic-based methods, as more experts are merged, the common information dominates the task-specific information, leading to inevitable rank collapse. To mitigate this issue, we introduce Subspace Boosting, which operates on the singular value decomposed task vector space and maintains task vector ranks. Subspace Boosting raises merging efficacy for up to 20 experts by large margins of more than 10% when evaluated on both vision and language benchmarks. Moreover, we propose employing Higher-Order Generalized Singular Value Decomposition to quantify task similarity, offering a new interpretable perspective on model merging. Code and models are available at https://github.com/ronskoro/Subspace-Boosting.
title Subspace-Boosted Model Merging
topic Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.16506