FINECAPTION: Compositional Image Captioning Focusing on Wherever You Want at Any Granularity

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hua, Hang, Liu, Qing, Zhang, Lingzhi, Shi, Jing, Zhang, Zhifei, Wang, Yilin, Zhang, Jianming, Luo, Jiebo
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912131095461888
author Hua, Hang
Liu, Qing
Zhang, Lingzhi
Shi, Jing
Zhang, Zhifei
Wang, Yilin
Zhang, Jianming
Luo, Jiebo
author_facet Hua, Hang
Liu, Qing
Zhang, Lingzhi
Shi, Jing
Zhang, Zhifei
Wang, Yilin
Zhang, Jianming
Luo, Jiebo
contents The advent of large Vision-Language Models (VLMs) has significantly advanced multimodal tasks, enabling more sophisticated and accurate reasoning across various applications, including image and video captioning, visual question answering, and cross-modal retrieval. Despite their superior capabilities, VLMs struggle with fine-grained image regional composition information perception. Specifically, they have difficulty accurately aligning the segmentation masks with the corresponding semantics and precisely describing the compositional aspects of the referred regions. However, compositionality - the ability to understand and generate novel combinations of known visual and textual components - is critical for facilitating coherent reasoning and understanding across modalities by VLMs. To address this issue, we propose FINECAPTION, a novel VLM that can recognize arbitrary masks as referential inputs and process high-resolution images for compositional image captioning at different granularity levels. To support this endeavor, we introduce COMPOSITIONCAP, a new dataset for multi-grained region compositional image captioning, which introduces the task of compositional attribute-aware regional image captioning. Empirical results demonstrate the effectiveness of our proposed model compared to other state-of-the-art VLMs. Additionally, we analyze the capabilities of current VLMs in recognizing various visual prompts for compositional region image captioning, highlighting areas for improvement in VLM design and training.
format Preprint
id arxiv_https___arxiv_org_abs_2411_15411
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle FINECAPTION: Compositional Image Captioning Focusing on Wherever You Want at Any Granularity
Hua, Hang
Liu, Qing
Zhang, Lingzhi
Shi, Jing
Zhang, Zhifei
Wang, Yilin
Zhang, Jianming
Luo, Jiebo
Computer Vision and Pattern Recognition
The advent of large Vision-Language Models (VLMs) has significantly advanced multimodal tasks, enabling more sophisticated and accurate reasoning across various applications, including image and video captioning, visual question answering, and cross-modal retrieval. Despite their superior capabilities, VLMs struggle with fine-grained image regional composition information perception. Specifically, they have difficulty accurately aligning the segmentation masks with the corresponding semantics and precisely describing the compositional aspects of the referred regions. However, compositionality - the ability to understand and generate novel combinations of known visual and textual components - is critical for facilitating coherent reasoning and understanding across modalities by VLMs. To address this issue, we propose FINECAPTION, a novel VLM that can recognize arbitrary masks as referential inputs and process high-resolution images for compositional image captioning at different granularity levels. To support this endeavor, we introduce COMPOSITIONCAP, a new dataset for multi-grained region compositional image captioning, which introduces the task of compositional attribute-aware regional image captioning. Empirical results demonstrate the effectiveness of our proposed model compared to other state-of-the-art VLMs. Additionally, we analyze the capabilities of current VLMs in recognizing various visual prompts for compositional region image captioning, highlighting areas for improvement in VLM design and training.
title FINECAPTION: Compositional Image Captioning Focusing on Wherever You Want at Any Granularity
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.15411