PhotoScout: Synthesis-Powered Multi-Modal Image Search

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Barnaby, Celeste, Chen, Qiaochu, Wang, Chenglong, Dillig, Isil
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913200351477760
author Barnaby, Celeste
Chen, Qiaochu
Wang, Chenglong
Dillig, Isil
author_facet Barnaby, Celeste
Chen, Qiaochu
Wang, Chenglong
Dillig, Isil
contents Due to the availability of increasingly large amounts of visual data, there is a growing need for tools that can help users find relevant images. While existing tools can perform image retrieval based on similarity or metadata, they fall short in scenarios that necessitate semantic reasoning about the content of the image. This paper explores a new multi-modal image search approach that allows users to conveniently specify and perform semantic image search tasks. With our tool, PhotoScout, the user interactively provides natural language descriptions, positive and negative examples, and object tags to specify their search tasks. Under the hood, PhotoScout is powered by a program synthesis engine that generates visual queries in a domain-specific language and executes the synthesized program to retrieve the desired images. In a study with 25 participants, we observed that PhotoScout allows users to perform image retrieval tasks more accurately and with less manual effort.
format Preprint
id arxiv_https___arxiv_org_abs_2401_10464
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle PhotoScout: Synthesis-Powered Multi-Modal Image Search
Barnaby, Celeste
Chen, Qiaochu
Wang, Chenglong
Dillig, Isil
Human-Computer Interaction
Due to the availability of increasingly large amounts of visual data, there is a growing need for tools that can help users find relevant images. While existing tools can perform image retrieval based on similarity or metadata, they fall short in scenarios that necessitate semantic reasoning about the content of the image. This paper explores a new multi-modal image search approach that allows users to conveniently specify and perform semantic image search tasks. With our tool, PhotoScout, the user interactively provides natural language descriptions, positive and negative examples, and object tags to specify their search tasks. Under the hood, PhotoScout is powered by a program synthesis engine that generates visual queries in a domain-specific language and executes the synthesized program to retrieve the desired images. In a study with 25 participants, we observed that PhotoScout allows users to perform image retrieval tasks more accurately and with less manual effort.
title PhotoScout: Synthesis-Powered Multi-Modal Image Search
topic Human-Computer Interaction
url https://arxiv.org/abs/2401.10464