Learning to Instruct for Visual Instruction Tuning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhou, Zhihan, Hong, Feng, Luo, Jiaan, Yao, Jiangchao, Li, Dongsheng, Han, Bo, Zhang, Ya, Wang, Yanfeng
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914087224475648
author Zhou, Zhihan
Hong, Feng
Luo, Jiaan
Yao, Jiangchao
Li, Dongsheng
Han, Bo
Zhang, Ya
Wang, Yanfeng
author_facet Zhou, Zhihan
Hong, Feng
Luo, Jiaan
Yao, Jiangchao
Li, Dongsheng
Han, Bo
Zhang, Ya
Wang, Yanfeng
contents We propose L2T, an advancement of visual instruction tuning (VIT). While VIT equips Multimodal LLMs (MLLMs) with promising multimodal capabilities, the current design choices for VIT often result in overfitting and shortcut learning, potentially degrading performance. This gap arises from an overemphasis on instruction-following abilities, while neglecting the proactive understanding of visual information. Inspired by this, L2T adopts a simple yet effective approach by incorporating the loss function into both the instruction and response sequences. It seamlessly expands the training data, and regularizes the MLLMs from overly relying on language priors. Based on this merit, L2T achieves a significant relative improvement of up to 9% on comprehensive multimodal benchmarks, requiring no additional training data and incurring negligible computational overhead. Surprisingly, L2T attains exceptional fundamental visual capabilities, yielding up to an 18% improvement in captioning performance, while simultaneously alleviating hallucination in MLLMs. Github code: https://github.com/Feng-Hong/L2T.
format Preprint
id arxiv_https___arxiv_org_abs_2503_22215
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning to Instruct for Visual Instruction Tuning
Zhou, Zhihan
Hong, Feng
Luo, Jiaan
Yao, Jiangchao
Li, Dongsheng
Han, Bo
Zhang, Ya
Wang, Yanfeng
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
We propose L2T, an advancement of visual instruction tuning (VIT). While VIT equips Multimodal LLMs (MLLMs) with promising multimodal capabilities, the current design choices for VIT often result in overfitting and shortcut learning, potentially degrading performance. This gap arises from an overemphasis on instruction-following abilities, while neglecting the proactive understanding of visual information. Inspired by this, L2T adopts a simple yet effective approach by incorporating the loss function into both the instruction and response sequences. It seamlessly expands the training data, and regularizes the MLLMs from overly relying on language priors. Based on this merit, L2T achieves a significant relative improvement of up to 9% on comprehensive multimodal benchmarks, requiring no additional training data and incurring negligible computational overhead. Surprisingly, L2T attains exceptional fundamental visual capabilities, yielding up to an 18% improvement in captioning performance, while simultaneously alleviating hallucination in MLLMs. Github code: https://github.com/Feng-Hong/L2T.
title Learning to Instruct for Visual Instruction Tuning
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2503.22215