How Well Can Vision Language Models See Image Details?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gou, Chenhui, Felemban, Abdulwahab, Khan, Faizan Farooq, Zhu, Deyao, Cai, Jianfei, Rezatofighi, Hamid, Elhoseiny, Mohamed
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916349958160384
author Gou, Chenhui
Felemban, Abdulwahab
Khan, Faizan Farooq
Zhu, Deyao
Cai, Jianfei
Rezatofighi, Hamid
Elhoseiny, Mohamed
author_facet Gou, Chenhui
Felemban, Abdulwahab
Khan, Faizan Farooq
Zhu, Deyao
Cai, Jianfei
Rezatofighi, Hamid
Elhoseiny, Mohamed
contents Large Language Model-based Vision-Language Models (LLM-based VLMs) have demonstrated impressive results in various vision-language understanding tasks. However, how well these VLMs can see image detail beyond the semantic level remains unclear. In our study, we introduce a pixel value prediction task (PVP) to explore "How Well Can Vision Language Models See Image Details?" and to assist VLMs in perceiving more details. Typically, these models comprise a frozen CLIP visual encoder, a large language model, and a connecting module. After fine-tuning VLMs on the PVP task, we find: 1) existing VLMs struggle to predict precise pixel values by only fine-tuning the connection module and LLM; and 2) prediction precision is significantly improved when the vision encoder is also adapted. Additionally, our research reveals that incorporating pixel value prediction as one of the VLM pre-training tasks and vision encoder adaptation markedly boosts VLM performance on downstream image-language understanding tasks requiring detailed image perception, such as referring image segmentation (with an average +10.19 cIoU improvement) and video game decision making (with average score improvements of +80.34 and +70.54 on two games, respectively).
format Preprint
id arxiv_https___arxiv_org_abs_2408_03940
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle How Well Can Vision Language Models See Image Details?
Gou, Chenhui
Felemban, Abdulwahab
Khan, Faizan Farooq
Zhu, Deyao
Cai, Jianfei
Rezatofighi, Hamid
Elhoseiny, Mohamed
Computer Vision and Pattern Recognition
Large Language Model-based Vision-Language Models (LLM-based VLMs) have demonstrated impressive results in various vision-language understanding tasks. However, how well these VLMs can see image detail beyond the semantic level remains unclear. In our study, we introduce a pixel value prediction task (PVP) to explore "How Well Can Vision Language Models See Image Details?" and to assist VLMs in perceiving more details. Typically, these models comprise a frozen CLIP visual encoder, a large language model, and a connecting module. After fine-tuning VLMs on the PVP task, we find: 1) existing VLMs struggle to predict precise pixel values by only fine-tuning the connection module and LLM; and 2) prediction precision is significantly improved when the vision encoder is also adapted. Additionally, our research reveals that incorporating pixel value prediction as one of the VLM pre-training tasks and vision encoder adaptation markedly boosts VLM performance on downstream image-language understanding tasks requiring detailed image perception, such as referring image segmentation (with an average +10.19 cIoU improvement) and video game decision making (with average score improvements of +80.34 and +70.54 on two games, respectively).
title How Well Can Vision Language Models See Image Details?
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2408.03940