ViPCap: Retrieval Text-Based Visual Prompts for Lightweight Image Captioning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Taewhan, Lee, Soeun, Kim, Si-Woo, Kim, Dong-Jin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915119166914560
author Kim, Taewhan
Lee, Soeun
Kim, Si-Woo
Kim, Dong-Jin
author_facet Kim, Taewhan
Lee, Soeun
Kim, Si-Woo
Kim, Dong-Jin
contents Recent lightweight image captioning models using retrieved data mainly focus on text prompts. However, previous works only utilize the retrieved text as text prompts, and the visual information relies only on the CLIP visual embedding. Because of this issue, there is a limitation that the image descriptions inherent in the prompt are not sufficiently reflected in the visual embedding space. To tackle this issue, we propose ViPCap, a novel retrieval text-based visual prompt for lightweight image captioning. ViPCap leverages the retrieved text with image information as visual prompts to enhance the ability of the model to capture relevant visual information. By mapping text prompts into the CLIP space and generating multiple randomized Gaussian distributions, our method leverages sampling to explore randomly augmented distributions and effectively retrieves the semantic features that contain image information. These retrieved features are integrated into the image and designated as the visual prompt, leading to performance improvements on the datasets such as COCO, Flickr30k, and NoCaps. Experimental results demonstrate that ViPCap significantly outperforms prior lightweight captioning models in efficiency and effectiveness, demonstrating the potential for a plug-and-play solution. The source code is available at https://github.com/taewhankim/VIPCAP.
format Preprint
id arxiv_https___arxiv_org_abs_2412_19289
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ViPCap: Retrieval Text-Based Visual Prompts for Lightweight Image Captioning
Kim, Taewhan
Lee, Soeun
Kim, Si-Woo
Kim, Dong-Jin
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Recent lightweight image captioning models using retrieved data mainly focus on text prompts. However, previous works only utilize the retrieved text as text prompts, and the visual information relies only on the CLIP visual embedding. Because of this issue, there is a limitation that the image descriptions inherent in the prompt are not sufficiently reflected in the visual embedding space. To tackle this issue, we propose ViPCap, a novel retrieval text-based visual prompt for lightweight image captioning. ViPCap leverages the retrieved text with image information as visual prompts to enhance the ability of the model to capture relevant visual information. By mapping text prompts into the CLIP space and generating multiple randomized Gaussian distributions, our method leverages sampling to explore randomly augmented distributions and effectively retrieves the semantic features that contain image information. These retrieved features are integrated into the image and designated as the visual prompt, leading to performance improvements on the datasets such as COCO, Flickr30k, and NoCaps. Experimental results demonstrate that ViPCap significantly outperforms prior lightweight captioning models in efficiency and effectiveness, demonstrating the potential for a plug-and-play solution. The source code is available at https://github.com/taewhankim/VIPCAP.
title ViPCap: Retrieval Text-Based Visual Prompts for Lightweight Image Captioning
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2412.19289