Do we Really Need Visual Instructions? Towards Visual Instruction-Free Fine-tuning for Large Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Zikang, Zhou, Kun, Zhao, Wayne Xin, Gao, Dawei, Li, Yaliang, Wen, Ji-Rong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929717907554304
author Liu, Zikang
Zhou, Kun
Zhao, Wayne Xin
Gao, Dawei
Li, Yaliang
Wen, Ji-Rong
author_facet Liu, Zikang
Zhou, Kun
Zhao, Wayne Xin
Gao, Dawei
Li, Yaliang
Wen, Ji-Rong
contents Visual instruction tuning has become the predominant technology in eliciting the multimodal task-solving capabilities of large vision-language models (LVLMs). Despite the success, as visual instructions require images as the input, it would leave the gap in inheriting the task-solving capabilities from the backbone LLMs, and make it costly to collect a large-scale dataset. To address it, we propose ViFT, a visual instruction-free fine-tuning framework for LVLMs. In ViFT, we only require the text-only instructions and image caption data during training, to separately learn the task-solving and visual perception abilities. During inference, we extract and combine the representations of the text and image inputs, for fusing the two abilities to fulfill multimodal tasks. Experimental results demonstrate that ViFT can achieve state-of-the-art performance on several visual reasoning and visual instruction following benchmarks, with rather less training data. Our code and data will be publicly released.
format Preprint
id arxiv_https___arxiv_org_abs_2502_11427
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Do we Really Need Visual Instructions? Towards Visual Instruction-Free Fine-tuning for Large Vision-Language Models
Liu, Zikang
Zhou, Kun
Zhao, Wayne Xin
Gao, Dawei
Li, Yaliang
Wen, Ji-Rong
Computation and Language
Computer Vision and Pattern Recognition
Visual instruction tuning has become the predominant technology in eliciting the multimodal task-solving capabilities of large vision-language models (LVLMs). Despite the success, as visual instructions require images as the input, it would leave the gap in inheriting the task-solving capabilities from the backbone LLMs, and make it costly to collect a large-scale dataset. To address it, we propose ViFT, a visual instruction-free fine-tuning framework for LVLMs. In ViFT, we only require the text-only instructions and image caption data during training, to separately learn the task-solving and visual perception abilities. During inference, we extract and combine the representations of the text and image inputs, for fusing the two abilities to fulfill multimodal tasks. Experimental results demonstrate that ViFT can achieve state-of-the-art performance on several visual reasoning and visual instruction following benchmarks, with rather less training data. Our code and data will be publicly released.
title Do we Really Need Visual Instructions? Towards Visual Instruction-Free Fine-tuning for Large Vision-Language Models
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2502.11427