Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Xiyao, Yang, Zhengyuan, Li, Linjie, Lu, Hongjin, Xu, Yuancheng, Lin, Chung-Ching, Lin, Kevin, Huang, Furong, Wang, Lijuan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909667640213504
author Wang, Xiyao
Yang, Zhengyuan
Li, Linjie
Lu, Hongjin
Xu, Yuancheng
Lin, Chung-Ching
Lin, Kevin
Huang, Furong
Wang, Lijuan
author_facet Wang, Xiyao
Yang, Zhengyuan
Li, Linjie
Lu, Hongjin
Xu, Yuancheng
Lin, Chung-Ching
Lin, Kevin
Huang, Furong
Wang, Lijuan
contents Despite significant advancements in vision-language models (VLMs), there lacks effective approaches to enhance response quality by scaling inference-time computation. This capability is known to be a core step towards the self-improving models in recent large language model studies. In this paper, we present Vision Value Model (VisVM) that can guide VLM inference-time search to generate responses with better visual comprehension. Specifically, VisVM not only evaluates the generated sentence quality in the current search step, but also anticipates the quality of subsequent sentences that may result from the current step, thus providing a long-term value. In this way, VisVM steers VLMs away from generating sentences prone to hallucinations or insufficient detail, thereby producing higher quality responses. Experimental results demonstrate that VisVM-guided search significantly enhances VLMs' ability to generate descriptive captions with richer visual details and fewer hallucinations, compared with greedy decoding and search methods with other visual reward signals. Furthermore, we find that self-training the model with the VisVM-guided captions improve VLM's performance across a wide range of multimodal benchmarks, indicating the potential for developing self-improving VLMs. Our value model and code are available at https://github.com/si0wang/VisVM.
format Preprint
id arxiv_https___arxiv_org_abs_2412_03704
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension
Wang, Xiyao
Yang, Zhengyuan
Li, Linjie
Lu, Hongjin
Xu, Yuancheng
Lin, Chung-Ching
Lin, Kevin
Huang, Furong
Wang, Lijuan
Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
Despite significant advancements in vision-language models (VLMs), there lacks effective approaches to enhance response quality by scaling inference-time computation. This capability is known to be a core step towards the self-improving models in recent large language model studies. In this paper, we present Vision Value Model (VisVM) that can guide VLM inference-time search to generate responses with better visual comprehension. Specifically, VisVM not only evaluates the generated sentence quality in the current search step, but also anticipates the quality of subsequent sentences that may result from the current step, thus providing a long-term value. In this way, VisVM steers VLMs away from generating sentences prone to hallucinations or insufficient detail, thereby producing higher quality responses. Experimental results demonstrate that VisVM-guided search significantly enhances VLMs' ability to generate descriptive captions with richer visual details and fewer hallucinations, compared with greedy decoding and search methods with other visual reward signals. Furthermore, we find that self-training the model with the VisVM-guided captions improve VLM's performance across a wide range of multimodal benchmarks, indicating the potential for developing self-improving VLMs. Our value model and code are available at https://github.com/si0wang/VisVM.
title Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension
topic Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
url https://arxiv.org/abs/2412.03704