Domain-Adaptive Model Merging Across Disconnected Modes

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Liu, Junming, Zhang, Yusen, Zhang, Rongchao, Zhu, Wenkai, Wu, Tian
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915937533296640
author Liu, Junming
Zhang, Yusen
Zhang, Rongchao
Zhu, Wenkai
Wu, Tian
author_facet Liu, Junming
Zhang, Yusen
Zhang, Rongchao
Zhu, Wenkai
Wu, Tian
contents Learning across domains is challenging when data cannot be centralized due to privacy or heterogeneity, which limits the ability to train a single comprehensive model. Model merging provides an appealing alternative by consolidating knowledge from multiple specialized models into one, avoiding data sharing and reducing retraining cost. In this work, we present DMM, a data-free model merging framework designed to handle highly divergent models. DMM proceeds in three steps. First, domain-specific models are trained independently. Second, models with high similarity are merged using standard techniques to ensure stability. Third, we synthesize pseudo-data from normalization statistics and distill knowledge from divergent models into the merged model through a lightweight refinement guided by these samples. This approach preserves rare but critical knowledge while maintaining stability. Extensive experiments on unimodal and multimodal benchmarks show that DMM achieves state-of-the-art performance over existing merging methods.
format Preprint
id arxiv_https___arxiv_org_abs_2603_05957
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Domain-Adaptive Model Merging Across Disconnected Modes
Liu, Junming
Zhang, Yusen
Zhang, Rongchao
Zhu, Wenkai
Wu, Tian
Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Learning across domains is challenging when data cannot be centralized due to privacy or heterogeneity, which limits the ability to train a single comprehensive model. Model merging provides an appealing alternative by consolidating knowledge from multiple specialized models into one, avoiding data sharing and reducing retraining cost. In this work, we present DMM, a data-free model merging framework designed to handle highly divergent models. DMM proceeds in three steps. First, domain-specific models are trained independently. Second, models with high similarity are merged using standard techniques to ensure stability. Third, we synthesize pseudo-data from normalization statistics and distill knowledge from divergent models into the merged model through a lightweight refinement guided by these samples. This approach preserves rare but critical knowledge while maintaining stability. Extensive experiments on unimodal and multimodal benchmarks show that DMM achieves state-of-the-art performance over existing merging methods.
title Domain-Adaptive Model Merging Across Disconnected Modes
topic Distributed, Parallel, and Cluster Computing
Artificial Intelligence
url https://arxiv.org/abs/2603.05957