Dynamic Embedding of Hierarchical Visual Features for Efficient Vision-Language Fine-Tuning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wei, Xinyu, Yang, Guoli, Zhou, Jialu, Yang, Mingyue, Li, Leqian, Zhang, Kedi, Qiu, Chunping
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918129919066112
author Wei, Xinyu
Yang, Guoli
Zhou, Jialu
Yang, Mingyue
Li, Leqian
Zhang, Kedi
Qiu, Chunping
author_facet Wei, Xinyu
Yang, Guoli
Zhou, Jialu
Yang, Mingyue
Li, Leqian
Zhang, Kedi
Qiu, Chunping
contents Large Vision-Language Models (LVLMs) commonly follow a paradigm that projects visual features and then concatenates them with text tokens to form a unified sequence input for Large Language Models (LLMs). However, this paradigm leads to a significant increase in the length of the input sequence, resulting in substantial computational overhead. Existing methods attempt to fuse visual information into the intermediate layers of LLMs, which alleviate the sequence length issue but often neglect the hierarchical semantic representations within the model and the fine-grained visual information available in the shallower visual encoding layers. To address this limitation, we propose DEHVF, an efficient vision-language fine-tuning method based on dynamic embedding and fusion of hierarchical visual features. Its core lies in leveraging the inherent hierarchical representation characteristics of visual encoders and language models. Through a lightweight hierarchical visual fuser, it dynamically selects and fuses hierarchical features corresponding to semantic granularity based on the internal representations of each layer in LLMs. The fused layer-related visual features are then projected and aligned before being directly embedded into the Feed-Forward Network (FFN) of the corresponding layer in LLMs. This approach not only avoids sequence expansion but also dynamically fuses multi-layer visual information. By fine-tuning only a small number of parameters, DEHVF achieves precise alignment and complementarity of cross-modal information at the same semantic granularity. We conducted experiments across various VL benchmarks, including visual question answering on ScienceQA and image captioning on COCO Captions. The results demonstrate that DEHVF achieves higher accuracy than existing parameter-efficient fine-tuning (PEFT) baselines while maintaining efficient training and inference.
format Preprint
id arxiv_https___arxiv_org_abs_2508_17638
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Dynamic Embedding of Hierarchical Visual Features for Efficient Vision-Language Fine-Tuning
Wei, Xinyu
Yang, Guoli
Zhou, Jialu
Yang, Mingyue
Li, Leqian
Zhang, Kedi
Qiu, Chunping
Computer Vision and Pattern Recognition
Computation and Language
Large Vision-Language Models (LVLMs) commonly follow a paradigm that projects visual features and then concatenates them with text tokens to form a unified sequence input for Large Language Models (LLMs). However, this paradigm leads to a significant increase in the length of the input sequence, resulting in substantial computational overhead. Existing methods attempt to fuse visual information into the intermediate layers of LLMs, which alleviate the sequence length issue but often neglect the hierarchical semantic representations within the model and the fine-grained visual information available in the shallower visual encoding layers. To address this limitation, we propose DEHVF, an efficient vision-language fine-tuning method based on dynamic embedding and fusion of hierarchical visual features. Its core lies in leveraging the inherent hierarchical representation characteristics of visual encoders and language models. Through a lightweight hierarchical visual fuser, it dynamically selects and fuses hierarchical features corresponding to semantic granularity based on the internal representations of each layer in LLMs. The fused layer-related visual features are then projected and aligned before being directly embedded into the Feed-Forward Network (FFN) of the corresponding layer in LLMs. This approach not only avoids sequence expansion but also dynamically fuses multi-layer visual information. By fine-tuning only a small number of parameters, DEHVF achieves precise alignment and complementarity of cross-modal information at the same semantic granularity. We conducted experiments across various VL benchmarks, including visual question answering on ScienceQA and image captioning on COCO Captions. The results demonstrate that DEHVF achieves higher accuracy than existing parameter-efficient fine-tuning (PEFT) baselines while maintaining efficient training and inference.
title Dynamic Embedding of Hierarchical Visual Features for Efficient Vision-Language Fine-Tuning
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2508.17638