Scalable Model Merging with Progressive Layer-wise Distillation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Jing, Li, Jiazheng, Zhang, Jingzhao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908381057384448
author Xu, Jing
Li, Jiazheng
Zhang, Jingzhao
author_facet Xu, Jing
Li, Jiazheng
Zhang, Jingzhao
contents Model merging offers an effective way to integrate the capabilities of multiple fine-tuned models. However, the performance degradation of the merged model remains a challenge, particularly when none or few data are available. This paper first highlights the necessity of domain-specific data for model merging by proving that data-agnostic algorithms can have arbitrarily bad worst-case performance. Building on this theoretical insight, we explore the relationship between model merging and distillation, introducing a novel few-shot merging algorithm, ProDistill (Progressive Layer-wise Distillation). Unlike common belief that layer wise training hurts performance, we show that layer-wise teacher-student distillation not only enhances the scalability but also improves model merging performance. We conduct extensive experiments to show that compared to existing few-shot merging methods, ProDistill achieves state-of-the-art performance, with up to 6.14% and 6.61% improvements in vision and NLU tasks. Furthermore, we extend the experiments to models with over 10B parameters, showcasing the exceptional scalability of ProDistill.
format Preprint
id arxiv_https___arxiv_org_abs_2502_12706
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Scalable Model Merging with Progressive Layer-wise Distillation
Xu, Jing
Li, Jiazheng
Zhang, Jingzhao
Machine Learning
Model merging offers an effective way to integrate the capabilities of multiple fine-tuned models. However, the performance degradation of the merged model remains a challenge, particularly when none or few data are available. This paper first highlights the necessity of domain-specific data for model merging by proving that data-agnostic algorithms can have arbitrarily bad worst-case performance. Building on this theoretical insight, we explore the relationship between model merging and distillation, introducing a novel few-shot merging algorithm, ProDistill (Progressive Layer-wise Distillation). Unlike common belief that layer wise training hurts performance, we show that layer-wise teacher-student distillation not only enhances the scalability but also improves model merging performance. We conduct extensive experiments to show that compared to existing few-shot merging methods, ProDistill achieves state-of-the-art performance, with up to 6.14% and 6.61% improvements in vision and NLU tasks. Furthermore, we extend the experiments to models with over 10B parameters, showcasing the exceptional scalability of ProDistill.
title Scalable Model Merging with Progressive Layer-wise Distillation
topic Machine Learning
url https://arxiv.org/abs/2502.12706