Mismatch Quest: Visual and Textual Feedback for Image-Text Misalignment

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Gordon, Brian, Bitton, Yonatan, Shafir, Yonatan, Garg, Roopal, Chen, Xi, Lischinski, Dani, Cohen-Or, Daniel, Szpektor, Idan
Format: Preprint
Publié: 2023
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917724092891136
author Gordon, Brian
Bitton, Yonatan
Shafir, Yonatan
Garg, Roopal
Chen, Xi
Lischinski, Dani
Cohen-Or, Daniel
Szpektor, Idan
author_facet Gordon, Brian
Bitton, Yonatan
Shafir, Yonatan
Garg, Roopal
Chen, Xi
Lischinski, Dani
Cohen-Or, Daniel
Szpektor, Idan
contents While existing image-text alignment models reach high quality binary assessments, they fall short of pinpointing the exact source of misalignment. In this paper, we present a method to provide detailed textual and visual explanation of detected misalignments between text-image pairs. We leverage large language models and visual grounding models to automatically construct a training set that holds plausible misaligned captions for a given image and corresponding textual explanations and visual indicators. We also publish a new human curated test set comprising ground-truth textual and visual misalignment annotations. Empirical results show that fine-tuning vision language models on our training set enables them to articulate misalignments and visually indicate them within images, outperforming strong baselines both on the binary alignment classification and the explanation generation tasks. Our method code and human curated test set are available at: https://mismatch-quest.github.io/
format Preprint
id arxiv_https___arxiv_org_abs_2312_03766
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Mismatch Quest: Visual and Textual Feedback for Image-Text Misalignment
Gordon, Brian
Bitton, Yonatan
Shafir, Yonatan
Garg, Roopal
Chen, Xi
Lischinski, Dani
Cohen-Or, Daniel
Szpektor, Idan
Computation and Language
Computer Vision and Pattern Recognition
While existing image-text alignment models reach high quality binary assessments, they fall short of pinpointing the exact source of misalignment. In this paper, we present a method to provide detailed textual and visual explanation of detected misalignments between text-image pairs. We leverage large language models and visual grounding models to automatically construct a training set that holds plausible misaligned captions for a given image and corresponding textual explanations and visual indicators. We also publish a new human curated test set comprising ground-truth textual and visual misalignment annotations. Empirical results show that fine-tuning vision language models on our training set enables them to articulate misalignments and visually indicate them within images, outperforming strong baselines both on the binary alignment classification and the explanation generation tasks. Our method code and human curated test set are available at: https://mismatch-quest.github.io/
title Mismatch Quest: Visual and Textual Feedback for Image-Text Misalignment
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2312.03766