The Role of Data Curation in Image Captioning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Li, Wenyan, Lotz, Jonas F., Qiu, Chen, Elliott, Desmond
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914663920304128
author Li, Wenyan
Lotz, Jonas F.
Qiu, Chen
Elliott, Desmond
author_facet Li, Wenyan
Lotz, Jonas F.
Qiu, Chen
Elliott, Desmond
contents Image captioning models are typically trained by treating all samples equally, neglecting to account for mismatched or otherwise difficult data points. In contrast, recent work has shown the effectiveness of training models by scheduling the data using curriculum learning strategies. This paper contributes to this direction by actively curating difficult samples in datasets without increasing the total number of samples. We explore the effect of using three data curation methods within the training process: complete removal of an sample, caption replacement, or image replacement via a text-to-image generation model. Experiments on the Flickr30K and COCO datasets with the BLIP and BEiT-3 models demonstrate that these curation methods do indeed yield improved image captioning models, underscoring their efficacy.
format Preprint
id arxiv_https___arxiv_org_abs_2305_03610
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle The Role of Data Curation in Image Captioning
Li, Wenyan
Lotz, Jonas F.
Qiu, Chen
Elliott, Desmond
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Image captioning models are typically trained by treating all samples equally, neglecting to account for mismatched or otherwise difficult data points. In contrast, recent work has shown the effectiveness of training models by scheduling the data using curriculum learning strategies. This paper contributes to this direction by actively curating difficult samples in datasets without increasing the total number of samples. We explore the effect of using three data curation methods within the training process: complete removal of an sample, caption replacement, or image replacement via a text-to-image generation model. Experiments on the Flickr30K and COCO datasets with the BLIP and BEiT-3 models demonstrate that these curation methods do indeed yield improved image captioning models, underscoring their efficacy.
title The Role of Data Curation in Image Captioning
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2305.03610