Exploring Diverse In-Context Configurations for Image Captioning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Xu, Wu, Yongliang, Yang, Mingzhuo, Chen, Haokun, Geng, Xin
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929218826272768
author Yang, Xu
Wu, Yongliang
Yang, Mingzhuo
Chen, Haokun
Geng, Xin
author_facet Yang, Xu
Wu, Yongliang
Yang, Mingzhuo
Chen, Haokun
Geng, Xin
contents After discovering that Language Models (LMs) can be good in-context few-shot learners, numerous strategies have been proposed to optimize in-context sequence configurations. Recently, researchers in Vision-Language (VL) domains also develop their few-shot learners, while they only use the simplest way, ie., randomly sampling, to configure in-context image-text pairs. In order to explore the effects of varying configurations on VL in-context learning, we devised four strategies for image selection and four for caption assignment to configure in-context image-text pairs for image captioning. Here Image Captioning is used as the case study since it can be seen as the visually-conditioned LM. Our comprehensive experiments yield two counter-intuitive but valuable insights, highlighting the distinct characteristics of VL in-context learning due to multi-modal synergy, as compared to the NLP case. Furthermore, in our exploration of optimal combination strategies, we observed an average performance enhancement of 20.9 of CIDEr scores compared to the baseline. The code is given in https://github.com/yongliang-wu/ExploreCfg.
format Preprint
id arxiv_https___arxiv_org_abs_2305_14800
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Exploring Diverse In-Context Configurations for Image Captioning
Yang, Xu
Wu, Yongliang
Yang, Mingzhuo
Chen, Haokun
Geng, Xin
Computer Vision and Pattern Recognition
After discovering that Language Models (LMs) can be good in-context few-shot learners, numerous strategies have been proposed to optimize in-context sequence configurations. Recently, researchers in Vision-Language (VL) domains also develop their few-shot learners, while they only use the simplest way, ie., randomly sampling, to configure in-context image-text pairs. In order to explore the effects of varying configurations on VL in-context learning, we devised four strategies for image selection and four for caption assignment to configure in-context image-text pairs for image captioning. Here Image Captioning is used as the case study since it can be seen as the visually-conditioned LM. Our comprehensive experiments yield two counter-intuitive but valuable insights, highlighting the distinct characteristics of VL in-context learning due to multi-modal synergy, as compared to the NLP case. Furthermore, in our exploration of optimal combination strategies, we observed an average performance enhancement of 20.9 of CIDEr scores compared to the baseline. The code is given in https://github.com/yongliang-wu/ExploreCfg.
title Exploring Diverse In-Context Configurations for Image Captioning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2305.14800