Graft: Integrating the Domain Knowledge via Efficient Parameter Synergy for MLLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dai, Yang, An, Jianxiang, Lin, Tianwei, He, Hongyang, Huang, Hongzhe, Zhang, Wenqiao, Lv, Zheqi, Tang, Siliang, Zhuang, Yueting
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913919807782912
author Dai, Yang
An, Jianxiang
Lin, Tianwei
He, Hongyang
Huang, Hongzhe
Zhang, Wenqiao
Lv, Zheqi
Tang, Siliang
Zhuang, Yueting
author_facet Dai, Yang
An, Jianxiang
Lin, Tianwei
He, Hongyang
Huang, Hongzhe
Zhang, Wenqiao
Lv, Zheqi
Tang, Siliang
Zhuang, Yueting
contents Multimodal Large Language Models (MLLMs) have achieved success across various domains. However, their applicability tends to degrade when confronted with different types of data inputs, especially for MLLMs that have been fine-tuned for specific tasks. Despite its importance, the study of knowledge sharing among domain-specific MLLMs--such as those trained for mathematics or code--remains largely underexplored. To address the fragmentation of knowledge across domain-specialized MLLMs, we propose a unified parameter integration framework that enables modular composition of expert capabilities. Our method is grounded in a novel Compatibility-Aware Parameter Splicing (CAPS) strategy, which leverages both local functional attribution and global information-theoretic signals to guide selective parameter fusion. By extending this mechanism to the low-rank adaptation layer granularity, we ensure efficient integration with minimal inference overhead. Furthermore, we introduce a domain compatibility scoring mechanism that quantifies inter-expert alignment at the activation level and correlates with downstream task utility. This principled fusion protocol allows the final model to synergize heterogeneous expertise while preserving structural modularity. Extensive evaluations across diverse multimodal benchmarks validate the effectiveness of our framework, offering a scalable path toward compositional, domain-adaptive MLLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2506_23940
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Graft: Integrating the Domain Knowledge via Efficient Parameter Synergy for MLLMs
Dai, Yang
An, Jianxiang
Lin, Tianwei
He, Hongyang
Huang, Hongzhe
Zhang, Wenqiao
Lv, Zheqi
Tang, Siliang
Zhuang, Yueting
Computation and Language
Multimodal Large Language Models (MLLMs) have achieved success across various domains. However, their applicability tends to degrade when confronted with different types of data inputs, especially for MLLMs that have been fine-tuned for specific tasks. Despite its importance, the study of knowledge sharing among domain-specific MLLMs--such as those trained for mathematics or code--remains largely underexplored. To address the fragmentation of knowledge across domain-specialized MLLMs, we propose a unified parameter integration framework that enables modular composition of expert capabilities. Our method is grounded in a novel Compatibility-Aware Parameter Splicing (CAPS) strategy, which leverages both local functional attribution and global information-theoretic signals to guide selective parameter fusion. By extending this mechanism to the low-rank adaptation layer granularity, we ensure efficient integration with minimal inference overhead. Furthermore, we introduce a domain compatibility scoring mechanism that quantifies inter-expert alignment at the activation level and correlates with downstream task utility. This principled fusion protocol allows the final model to synergize heterogeneous expertise while preserving structural modularity. Extensive evaluations across diverse multimodal benchmarks validate the effectiveness of our framework, offering a scalable path toward compositional, domain-adaptive MLLMs.
title Graft: Integrating the Domain Knowledge via Efficient Parameter Synergy for MLLMs
topic Computation and Language
url https://arxiv.org/abs/2506.23940