StableSemantics: A Synthetic Language-Vision Dataset of Semantic Representations in Naturalistic Images

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zawar, Rushikesh, Dewan, Shaurya, Luo, Andrew F., Henderson, Margaret M., Tarr, Michael J., Wehbe, Leila
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929392111845376
author Zawar, Rushikesh
Dewan, Shaurya
Luo, Andrew F.
Henderson, Margaret M.
Tarr, Michael J.
Wehbe, Leila
author_facet Zawar, Rushikesh
Dewan, Shaurya
Luo, Andrew F.
Henderson, Margaret M.
Tarr, Michael J.
Wehbe, Leila
contents Understanding the semantics of visual scenes is a fundamental challenge in Computer Vision. A key aspect of this challenge is that objects sharing similar semantic meanings or functions can exhibit striking visual differences, making accurate identification and categorization difficult. Recent advancements in text-to-image frameworks have led to models that implicitly capture natural scene statistics. These frameworks account for the visual variability of objects, as well as complex object co-occurrences and sources of noise such as diverse lighting conditions. By leveraging large-scale datasets and cross-attention conditioning, these models generate detailed and contextually rich scene representations. This capability opens new avenues for improving object recognition and scene understanding in varied and challenging environments. Our work presents StableSemantics, a dataset comprising 224 thousand human-curated prompts, processed natural language captions, over 2 million synthetic images, and 10 million attention maps corresponding to individual noun chunks. We explicitly leverage human-generated prompts that correspond to visually interesting stable diffusion generations, provide 10 generations per phrase, and extract cross-attention maps for each image. We explore the semantic distribution of generated images, examine the distribution of objects within images, and benchmark captioning and open vocabulary segmentation methods on our data. To the best of our knowledge, we are the first to release a diffusion dataset with semantic attributions. We expect our proposed dataset to catalyze advances in visual semantic understanding and provide a foundation for developing more sophisticated and effective visual models. Website: https://stablesemantics.github.io/StableSemantics
format Preprint
id arxiv_https___arxiv_org_abs_2406_13735
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle StableSemantics: A Synthetic Language-Vision Dataset of Semantic Representations in Naturalistic Images
Zawar, Rushikesh
Dewan, Shaurya
Luo, Andrew F.
Henderson, Margaret M.
Tarr, Michael J.
Wehbe, Leila
Computer Vision and Pattern Recognition
Machine Learning
Understanding the semantics of visual scenes is a fundamental challenge in Computer Vision. A key aspect of this challenge is that objects sharing similar semantic meanings or functions can exhibit striking visual differences, making accurate identification and categorization difficult. Recent advancements in text-to-image frameworks have led to models that implicitly capture natural scene statistics. These frameworks account for the visual variability of objects, as well as complex object co-occurrences and sources of noise such as diverse lighting conditions. By leveraging large-scale datasets and cross-attention conditioning, these models generate detailed and contextually rich scene representations. This capability opens new avenues for improving object recognition and scene understanding in varied and challenging environments. Our work presents StableSemantics, a dataset comprising 224 thousand human-curated prompts, processed natural language captions, over 2 million synthetic images, and 10 million attention maps corresponding to individual noun chunks. We explicitly leverage human-generated prompts that correspond to visually interesting stable diffusion generations, provide 10 generations per phrase, and extract cross-attention maps for each image. We explore the semantic distribution of generated images, examine the distribution of objects within images, and benchmark captioning and open vocabulary segmentation methods on our data. To the best of our knowledge, we are the first to release a diffusion dataset with semantic attributions. We expect our proposed dataset to catalyze advances in visual semantic understanding and provide a foundation for developing more sophisticated and effective visual models. Website: https://stablesemantics.github.io/StableSemantics
title StableSemantics: A Synthetic Language-Vision Dataset of Semantic Representations in Naturalistic Images
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2406.13735