Saved in:
Bibliographic Details
Main Authors: Xu, Zhiyu, Wang, Lean, Liu, Yuanxin, Li, Lei, Zhou, Hao, Meng, Fandong, Zhou, Jie, Sun, Xu
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2605.19523
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910235855159296
author Xu, Zhiyu
Wang, Lean
Liu, Yuanxin
Li, Lei
Zhou, Hao
Meng, Fandong
Zhou, Jie
Sun, Xu
author_facet Xu, Zhiyu
Wang, Lean
Liu, Yuanxin
Li, Lei
Zhou, Hao
Meng, Fandong
Zhou, Jie
Sun, Xu
contents Vision-Language Models (VLMs) have demonstrated remarkable proficiency in general multi-modal understanding; yet they struggle to efficiently acquire continually evolving domain-specific skills. Conventional approaches to enhancing VLM capabilities, such as Supervised Fine-Tuning (SFT), require extensive dataset curation and substantial computational resources. Model merging has emerged as an efficient alternative that enables the transfer of domain-specific expertise from Large Language Models (LLMs) to VLMs without incurring additional training data requirements or significant computational overhead. Unlike conventional merging of homogeneous LLMs, which mainly aggregates existing capabilities, cross-modal skill injection aims to induce emergent cross-modal capabilities by integrating a domain-expert LLM into a VLM. However, existing research lacks a systematic analysis of the applicability and methodology of cross-modal skill injection. In this study, we investigate cross-modal skill injection across three main aspects: scenarios, methods, and hyperparameters. For scenarios, we find that cross-modal skill injection generally performs well in instruction-following and cross-lingual settings, yet struggles with mathematical reasoning. For methods, we find that classic approaches such as TA and DARE consistently achieve superior performance over alternative merging methods. We also provide a systematic and quantitative analysis of the hyperparameter tuning that these classic methods critically depend on.
format Preprint
id arxiv_https___arxiv_org_abs_2605_19523
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Investigating Cross-Modal Skill Injection: Scenarios, Methods, and Hyperparameters
Xu, Zhiyu
Wang, Lean
Liu, Yuanxin
Li, Lei
Zhou, Hao
Meng, Fandong
Zhou, Jie
Sun, Xu
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Vision-Language Models (VLMs) have demonstrated remarkable proficiency in general multi-modal understanding; yet they struggle to efficiently acquire continually evolving domain-specific skills. Conventional approaches to enhancing VLM capabilities, such as Supervised Fine-Tuning (SFT), require extensive dataset curation and substantial computational resources. Model merging has emerged as an efficient alternative that enables the transfer of domain-specific expertise from Large Language Models (LLMs) to VLMs without incurring additional training data requirements or significant computational overhead. Unlike conventional merging of homogeneous LLMs, which mainly aggregates existing capabilities, cross-modal skill injection aims to induce emergent cross-modal capabilities by integrating a domain-expert LLM into a VLM. However, existing research lacks a systematic analysis of the applicability and methodology of cross-modal skill injection. In this study, we investigate cross-modal skill injection across three main aspects: scenarios, methods, and hyperparameters. For scenarios, we find that cross-modal skill injection generally performs well in instruction-following and cross-lingual settings, yet struggles with mathematical reasoning. For methods, we find that classic approaches such as TA and DARE consistently achieve superior performance over alternative merging methods. We also provide a systematic and quantitative analysis of the hyperparameter tuning that these classic methods critically depend on.
title Investigating Cross-Modal Skill Injection: Scenarios, Methods, and Hyperparameters
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.19523