Describing Differences in Image Sets with Natural Language

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Dunlap, Lisa, Zhang, Yuhui, Wang, Xiaohan, Zhong, Ruiqi, Darrell, Trevor, Steinhardt, Jacob, Gonzalez, Joseph E., Yeung-Levy, Serena
Format: Preprint
Publié: 2023
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911856302489600
author Dunlap, Lisa
Zhang, Yuhui
Wang, Xiaohan
Zhong, Ruiqi
Darrell, Trevor
Steinhardt, Jacob
Gonzalez, Joseph E.
Yeung-Levy, Serena
author_facet Dunlap, Lisa
Zhang, Yuhui
Wang, Xiaohan
Zhong, Ruiqi
Darrell, Trevor
Steinhardt, Jacob
Gonzalez, Joseph E.
Yeung-Levy, Serena
contents How do two sets of images differ? Discerning set-level differences is crucial for understanding model behaviors and analyzing datasets, yet manually sifting through thousands of images is impractical. To aid in this discovery process, we explore the task of automatically describing the differences between two $\textbf{sets}$ of images, which we term Set Difference Captioning. This task takes in image sets $D_A$ and $D_B$, and outputs a description that is more often true on $D_A$ than $D_B$. We outline a two-stage approach that first proposes candidate difference descriptions from image sets and then re-ranks the candidates by checking how well they can differentiate the two sets. We introduce VisDiff, which first captions the images and prompts a language model to propose candidate descriptions, then re-ranks these descriptions using CLIP. To evaluate VisDiff, we collect VisDiffBench, a dataset with 187 paired image sets with ground truth difference descriptions. We apply VisDiff to various domains, such as comparing datasets (e.g., ImageNet vs. ImageNetV2), comparing classification models (e.g., zero-shot CLIP vs. supervised ResNet), summarizing model failure modes (supervised ResNet), characterizing differences between generative models (e.g., StableDiffusionV1 and V2), and discovering what makes images memorable. Using VisDiff, we are able to find interesting and previously unknown differences in datasets and models, demonstrating its utility in revealing nuanced insights.
format Preprint
id arxiv_https___arxiv_org_abs_2312_02974
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Describing Differences in Image Sets with Natural Language
Dunlap, Lisa
Zhang, Yuhui
Wang, Xiaohan
Zhong, Ruiqi
Darrell, Trevor
Steinhardt, Jacob
Gonzalez, Joseph E.
Yeung-Levy, Serena
Computer Vision and Pattern Recognition
Computation and Language
Computers and Society
Machine Learning
How do two sets of images differ? Discerning set-level differences is crucial for understanding model behaviors and analyzing datasets, yet manually sifting through thousands of images is impractical. To aid in this discovery process, we explore the task of automatically describing the differences between two $\textbf{sets}$ of images, which we term Set Difference Captioning. This task takes in image sets $D_A$ and $D_B$, and outputs a description that is more often true on $D_A$ than $D_B$. We outline a two-stage approach that first proposes candidate difference descriptions from image sets and then re-ranks the candidates by checking how well they can differentiate the two sets. We introduce VisDiff, which first captions the images and prompts a language model to propose candidate descriptions, then re-ranks these descriptions using CLIP. To evaluate VisDiff, we collect VisDiffBench, a dataset with 187 paired image sets with ground truth difference descriptions. We apply VisDiff to various domains, such as comparing datasets (e.g., ImageNet vs. ImageNetV2), comparing classification models (e.g., zero-shot CLIP vs. supervised ResNet), summarizing model failure modes (supervised ResNet), characterizing differences between generative models (e.g., StableDiffusionV1 and V2), and discovering what makes images memorable. Using VisDiff, we are able to find interesting and previously unknown differences in datasets and models, demonstrating its utility in revealing nuanced insights.
title Describing Differences in Image Sets with Natural Language
topic Computer Vision and Pattern Recognition
Computation and Language
Computers and Society
Machine Learning
url https://arxiv.org/abs/2312.02974