Toward Interactive Regional Understanding in Vision-Large Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Lee, Jungbeom, Chun, Sanghyuk, Yun, Sangdoo
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916179804684288
author Lee, Jungbeom
Chun, Sanghyuk
Yun, Sangdoo
author_facet Lee, Jungbeom
Chun, Sanghyuk
Yun, Sangdoo
contents Recent Vision-Language Pre-training (VLP) models have demonstrated significant advancements. Nevertheless, these models heavily rely on image-text pairs that capture only coarse and global information of an image, leading to a limitation in their regional understanding ability. In this work, we introduce \textbf{RegionVLM}, equipped with explicit regional modeling capabilities, allowing them to understand user-indicated image regions. To achieve this, we design a simple yet innovative architecture, requiring no modifications to the model architecture or objective function. Additionally, we leverage a dataset that contains a novel source of information, namely Localized Narratives, which has been overlooked in previous VLP research. Our experiments demonstrate that our single generalist model not only achieves an interactive dialogue system but also exhibits superior performance on various zero-shot region understanding tasks, without compromising its ability for global image understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2403_18260
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Toward Interactive Regional Understanding in Vision-Large Language Models
Lee, Jungbeom
Chun, Sanghyuk
Yun, Sangdoo
Computer Vision and Pattern Recognition
Computation and Language
Recent Vision-Language Pre-training (VLP) models have demonstrated significant advancements. Nevertheless, these models heavily rely on image-text pairs that capture only coarse and global information of an image, leading to a limitation in their regional understanding ability. In this work, we introduce \textbf{RegionVLM}, equipped with explicit regional modeling capabilities, allowing them to understand user-indicated image regions. To achieve this, we design a simple yet innovative architecture, requiring no modifications to the model architecture or objective function. Additionally, we leverage a dataset that contains a novel source of information, namely Localized Narratives, which has been overlooked in previous VLP research. Our experiments demonstrate that our single generalist model not only achieves an interactive dialogue system but also exhibits superior performance on various zero-shot region understanding tasks, without compromising its ability for global image understanding.
title Toward Interactive Regional Understanding in Vision-Large Language Models
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2403.18260