Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Shiqi, Zhang, Jinghan, Zhu, Tongyao, Liu, Wei, Gao, Siyang, Xiong, Miao, Li, Manling, He, Junxian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909689758875648
author Chen, Shiqi
Zhang, Jinghan
Zhu, Tongyao
Liu, Wei
Gao, Siyang
Xiong, Miao
Li, Manling
He, Junxian
author_facet Chen, Shiqi
Zhang, Jinghan
Zhu, Tongyao
Liu, Wei
Gao, Siyang
Xiong, Miao
Li, Manling
He, Junxian
contents Vision-Language Models (VLMs) combine visual perception with the general capabilities, such as reasoning, of Large Language Models (LLMs). However, the mechanisms by which these two abilities can be combined and contribute remain poorly understood. In this work, we explore to compose perception and reasoning through model merging that connects parameters of different models. Unlike previous works that often focus on merging models of the same kind, we propose merging models across modalities, enabling the incorporation of the reasoning capabilities of LLMs into VLMs. Through extensive experiments, we demonstrate that model merging offers a successful pathway to transfer reasoning abilities from LLMs to VLMs in a training-free manner. Moreover, we utilize the merged models to understand the internal mechanism of perception and reasoning and how merging affects it. We find that perception capabilities are predominantly encoded in the early layers of the model, whereas reasoning is largely facilitated by the middle-to-late layers. After merging, we observe that all layers begin to contribute to reasoning, whereas the distribution of perception abilities across layers remains largely unchanged. These observations shed light on the potential of model merging as a tool for multimodal integration and interpretation.
format Preprint
id arxiv_https___arxiv_org_abs_2505_05464
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging
Chen, Shiqi
Zhang, Jinghan
Zhu, Tongyao
Liu, Wei
Gao, Siyang
Xiong, Miao
Li, Manling
He, Junxian
Computation and Language
Vision-Language Models (VLMs) combine visual perception with the general capabilities, such as reasoning, of Large Language Models (LLMs). However, the mechanisms by which these two abilities can be combined and contribute remain poorly understood. In this work, we explore to compose perception and reasoning through model merging that connects parameters of different models. Unlike previous works that often focus on merging models of the same kind, we propose merging models across modalities, enabling the incorporation of the reasoning capabilities of LLMs into VLMs. Through extensive experiments, we demonstrate that model merging offers a successful pathway to transfer reasoning abilities from LLMs to VLMs in a training-free manner. Moreover, we utilize the merged models to understand the internal mechanism of perception and reasoning and how merging affects it. We find that perception capabilities are predominantly encoded in the early layers of the model, whereas reasoning is largely facilitated by the middle-to-late layers. After merging, we observe that all layers begin to contribute to reasoning, whereas the distribution of perception abilities across layers remains largely unchanged. These observations shed light on the potential of model merging as a tool for multimodal integration and interpretation.
title Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging
topic Computation and Language
url https://arxiv.org/abs/2505.05464