Multi-Layer Visual Feature Fusion in Multimodal LLMs: Methods, Analysis, and Best Practices

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lin, Junyan, Chen, Haoran, Fan, Yue, Fan, Yingqi, Jin, Xin, Su, Hui, Fu, Jinlan, Shen, Xiaoyu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910865122394112
author Lin, Junyan
Chen, Haoran
Fan, Yue
Fan, Yingqi
Jin, Xin
Su, Hui
Fu, Jinlan
Shen, Xiaoyu
author_facet Lin, Junyan
Chen, Haoran
Fan, Yue
Fan, Yingqi
Jin, Xin
Su, Hui
Fu, Jinlan
Shen, Xiaoyu
contents Multimodal Large Language Models (MLLMs) have made significant advancements in recent years, with visual features playing an increasingly critical role in enhancing model performance. However, the integration of multi-layer visual features in MLLMs remains underexplored, particularly with regard to optimal layer selection and fusion strategies. Existing methods often rely on arbitrary design choices, leading to suboptimal outcomes. In this paper, we systematically investigate two core aspects of multi-layer visual feature fusion: (1) selecting the most effective visual layers and (2) identifying the best fusion approach with the language model. Our experiments reveal that while combining visual features from multiple stages improves generalization, incorporating additional features from the same stage typically leads to diminished performance. Furthermore, we find that direct fusion of multi-layer visual features at the input stage consistently yields superior and more stable performance across various configurations. We make all our code publicly available: https://github.com/EIT-NLP/Layer_Select_Fuse_for_MLLM.
format Preprint
id arxiv_https___arxiv_org_abs_2503_06063
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multi-Layer Visual Feature Fusion in Multimodal LLMs: Methods, Analysis, and Best Practices
Lin, Junyan
Chen, Haoran
Fan, Yue
Fan, Yingqi
Jin, Xin
Su, Hui
Fu, Jinlan
Shen, Xiaoyu
Computer Vision and Pattern Recognition
Multimodal Large Language Models (MLLMs) have made significant advancements in recent years, with visual features playing an increasingly critical role in enhancing model performance. However, the integration of multi-layer visual features in MLLMs remains underexplored, particularly with regard to optimal layer selection and fusion strategies. Existing methods often rely on arbitrary design choices, leading to suboptimal outcomes. In this paper, we systematically investigate two core aspects of multi-layer visual feature fusion: (1) selecting the most effective visual layers and (2) identifying the best fusion approach with the language model. Our experiments reveal that while combining visual features from multiple stages improves generalization, incorporating additional features from the same stage typically leads to diminished performance. Furthermore, we find that direct fusion of multi-layer visual features at the input stage consistently yields superior and more stable performance across various configurations. We make all our code publicly available: https://github.com/EIT-NLP/Layer_Select_Fuse_for_MLLM.
title Multi-Layer Visual Feature Fusion in Multimodal LLMs: Methods, Analysis, and Best Practices
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.06063