ICC: Quantifying Image Caption Concreteness for Multimodal Dataset Curation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yanuka, Moran, Alper, Morris, Averbuch-Elor, Hadar, Giryes, Raja
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912533494890496
author Yanuka, Moran
Alper, Morris
Averbuch-Elor, Hadar
Giryes, Raja
author_facet Yanuka, Moran
Alper, Morris
Averbuch-Elor, Hadar
Giryes, Raja
contents Web-scale training on paired text-image data is becoming increasingly central to multimodal learning, but is challenged by the highly noisy nature of datasets in the wild. Standard data filtering approaches succeed in removing mismatched text-image pairs, but permit semantically related but highly abstract or subjective text. These approaches lack the fine-grained ability to isolate the most concrete samples that provide the strongest signal for learning in a noisy dataset. In this work, we propose a new metric, image caption concreteness, that evaluates caption text without an image reference to measure its concreteness and relevancy for use in multimodal learning. Our approach leverages strong foundation models for measuring visual-semantic information loss in multimodal representations. We demonstrate that this strongly correlates with human evaluation of concreteness in both single-word and sentence-level texts. Moreover, we show that curation using ICC complements existing approaches: It succeeds in selecting the highest quality samples from multimodal web-scale datasets to allow for efficient training in resource-constrained settings.
format Preprint
id arxiv_https___arxiv_org_abs_2403_01306
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ICC: Quantifying Image Caption Concreteness for Multimodal Dataset Curation
Yanuka, Moran
Alper, Morris
Averbuch-Elor, Hadar
Giryes, Raja
Machine Learning
Computer Vision and Pattern Recognition
Web-scale training on paired text-image data is becoming increasingly central to multimodal learning, but is challenged by the highly noisy nature of datasets in the wild. Standard data filtering approaches succeed in removing mismatched text-image pairs, but permit semantically related but highly abstract or subjective text. These approaches lack the fine-grained ability to isolate the most concrete samples that provide the strongest signal for learning in a noisy dataset. In this work, we propose a new metric, image caption concreteness, that evaluates caption text without an image reference to measure its concreteness and relevancy for use in multimodal learning. Our approach leverages strong foundation models for measuring visual-semantic information loss in multimodal representations. We demonstrate that this strongly correlates with human evaluation of concreteness in both single-word and sentence-level texts. Moreover, we show that curation using ICC complements existing approaches: It succeeds in selecting the highest quality samples from multimodal web-scale datasets to allow for efficient training in resource-constrained settings.
title ICC: Quantifying Image Caption Concreteness for Multimodal Dataset Curation
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.01306