OptMerge: Unifying Multimodal LLM Capabilities and Modalities via Model Merging

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wei, Yongxian, Cheng, Runxi, Jin, Weike, Yang, Enneng, Shen, Li, Hou, Lu, Du, Sinan, Yuan, Chun, Cao, Xiaochun, Tao, Dacheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911480428888064
author Wei, Yongxian
Cheng, Runxi
Jin, Weike
Yang, Enneng
Shen, Li
Hou, Lu
Du, Sinan
Yuan, Chun
Cao, Xiaochun
Tao, Dacheng
author_facet Wei, Yongxian
Cheng, Runxi
Jin, Weike
Yang, Enneng
Shen, Li
Hou, Lu
Du, Sinan
Yuan, Chun
Cao, Xiaochun
Tao, Dacheng
contents Foundation models update slowly due to resource-intensive training, whereas domain-specific models evolve rapidly between releases. Model merging seeks to combine multiple expert models into a single, more capable model, reducing storage and serving costs while supporting decentralized development. Despite its potential, previous studies have primarily focused on merging visual classification models or Large Language Models (LLMs) for code and math tasks. Recently, Multimodal LLMs (MLLMs) that extend LLMs through large-scale multimodal training have gained traction. However, there lacks a benchmark for model merging research that clearly divides the tasks for MLLM training and evaluation. In this paper, $\textbf{(i)}$ we introduce a model merging benchmark for MLLMs, which includes multiple tasks such as VQA, Geometry, Chart, OCR, and Grounding, studying both LoRA and full fine-tuning models. Moreover, we explore how model merging can combine different modalities (e.g., vision-language, audio-language, and video-language models), moving toward the Omni-language model. $\textbf{(ii)}$ We implement 10 model merging algorithms on the benchmark. Furthermore, we propose a novel method that removes noise from task vectors and robustly optimizes the merged vector based on a loss defined over task vector interactions, achieving an average performance gain of 2.48%. $\textbf{(iii)}$ We find that model merging offers a promising way for building improved MLLMs without requiring training data. Our results also demonstrate that the complementarity among multiple modalities outperforms individual modalities.
format Preprint
id arxiv_https___arxiv_org_abs_2505_19892
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OptMerge: Unifying Multimodal LLM Capabilities and Modalities via Model Merging
Wei, Yongxian
Cheng, Runxi
Jin, Weike
Yang, Enneng
Shen, Li
Hou, Lu
Du, Sinan
Yuan, Chun
Cao, Xiaochun
Tao, Dacheng
Artificial Intelligence
Foundation models update slowly due to resource-intensive training, whereas domain-specific models evolve rapidly between releases. Model merging seeks to combine multiple expert models into a single, more capable model, reducing storage and serving costs while supporting decentralized development. Despite its potential, previous studies have primarily focused on merging visual classification models or Large Language Models (LLMs) for code and math tasks. Recently, Multimodal LLMs (MLLMs) that extend LLMs through large-scale multimodal training have gained traction. However, there lacks a benchmark for model merging research that clearly divides the tasks for MLLM training and evaluation. In this paper, $\textbf{(i)}$ we introduce a model merging benchmark for MLLMs, which includes multiple tasks such as VQA, Geometry, Chart, OCR, and Grounding, studying both LoRA and full fine-tuning models. Moreover, we explore how model merging can combine different modalities (e.g., vision-language, audio-language, and video-language models), moving toward the Omni-language model. $\textbf{(ii)}$ We implement 10 model merging algorithms on the benchmark. Furthermore, we propose a novel method that removes noise from task vectors and robustly optimizes the merged vector based on a loss defined over task vector interactions, achieving an average performance gain of 2.48%. $\textbf{(iii)}$ We find that model merging offers a promising way for building improved MLLMs without requiring training data. Our results also demonstrate that the complementarity among multiple modalities outperforms individual modalities.
title OptMerge: Unifying Multimodal LLM Capabilities and Modalities via Model Merging
topic Artificial Intelligence
url https://arxiv.org/abs/2505.19892