Compositional Image Decomposition with Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Su, Jocelin, Liu, Nan, Wang, Yanbo, Tenenbaum, Joshua B., Du, Yilun
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909232560865280
author Su, Jocelin
Liu, Nan
Wang, Yanbo
Tenenbaum, Joshua B.
Du, Yilun
author_facet Su, Jocelin
Liu, Nan
Wang, Yanbo
Tenenbaum, Joshua B.
Du, Yilun
contents Given an image of a natural scene, we are able to quickly decompose it into a set of components such as objects, lighting, shadows, and foreground. We can then envision a scene where we combine certain components with those from other images, for instance a set of objects from our bedroom and animals from a zoo under the lighting conditions of a forest, even if we have never encountered such a scene before. In this paper, we present a method to decompose an image into such compositional components. Our approach, Decomp Diffusion, is an unsupervised method which, when given a single image, infers a set of different components in the image, each represented by a diffusion model. We demonstrate how components can capture different factors of the scene, ranging from global scene descriptors like shadows or facial expression to local scene descriptors like constituent objects. We further illustrate how inferred factors can be flexibly composed, even with factors inferred from other models, to generate a variety of scenes sharply different than those seen in training time. Website and code at https://energy-based-model.github.io/decomp-diffusion.
format Preprint
id arxiv_https___arxiv_org_abs_2406_19298
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Compositional Image Decomposition with Diffusion Models
Su, Jocelin
Liu, Nan
Wang, Yanbo
Tenenbaum, Joshua B.
Du, Yilun
Computer Vision and Pattern Recognition
Machine Learning
Given an image of a natural scene, we are able to quickly decompose it into a set of components such as objects, lighting, shadows, and foreground. We can then envision a scene where we combine certain components with those from other images, for instance a set of objects from our bedroom and animals from a zoo under the lighting conditions of a forest, even if we have never encountered such a scene before. In this paper, we present a method to decompose an image into such compositional components. Our approach, Decomp Diffusion, is an unsupervised method which, when given a single image, infers a set of different components in the image, each represented by a diffusion model. We demonstrate how components can capture different factors of the scene, ranging from global scene descriptors like shadows or facial expression to local scene descriptors like constituent objects. We further illustrate how inferred factors can be flexibly composed, even with factors inferred from other models, to generate a variety of scenes sharply different than those seen in training time. Website and code at https://energy-based-model.github.io/decomp-diffusion.
title Compositional Image Decomposition with Diffusion Models
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2406.19298