ZeroComp: Zero-shot Object Compositing from Image Intrinsics via Diffusion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Zitian, Fortier-Chouinard, Frédéric, Garon, Mathieu, Bhattad, Anand, Lalonde, Jean-François
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912182585786368
author Zhang, Zitian
Fortier-Chouinard, Frédéric
Garon, Mathieu
Bhattad, Anand
Lalonde, Jean-François
author_facet Zhang, Zitian
Fortier-Chouinard, Frédéric
Garon, Mathieu
Bhattad, Anand
Lalonde, Jean-François
contents We present ZeroComp, an effective zero-shot 3D object compositing approach that does not require paired composite-scene images during training. Our method leverages ControlNet to condition from intrinsic images and combines it with a Stable Diffusion model to utilize its scene priors, together operating as an effective rendering engine. During training, ZeroComp uses intrinsic images based on geometry, albedo, and masked shading, all without the need for paired images of scenes with and without composite objects. Once trained, it seamlessly integrates virtual 3D objects into scenes, adjusting shading to create realistic composites. We developed a high-quality evaluation dataset and demonstrate that ZeroComp outperforms methods using explicit lighting estimations and generative techniques in quantitative and human perception benchmarks. Additionally, ZeroComp extends to real and outdoor image compositing, even when trained solely on synthetic indoor data, showcasing its effectiveness in image compositing.
format Preprint
id arxiv_https___arxiv_org_abs_2410_08168
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ZeroComp: Zero-shot Object Compositing from Image Intrinsics via Diffusion
Zhang, Zitian
Fortier-Chouinard, Frédéric
Garon, Mathieu
Bhattad, Anand
Lalonde, Jean-François
Computer Vision and Pattern Recognition
We present ZeroComp, an effective zero-shot 3D object compositing approach that does not require paired composite-scene images during training. Our method leverages ControlNet to condition from intrinsic images and combines it with a Stable Diffusion model to utilize its scene priors, together operating as an effective rendering engine. During training, ZeroComp uses intrinsic images based on geometry, albedo, and masked shading, all without the need for paired images of scenes with and without composite objects. Once trained, it seamlessly integrates virtual 3D objects into scenes, adjusting shading to create realistic composites. We developed a high-quality evaluation dataset and demonstrate that ZeroComp outperforms methods using explicit lighting estimations and generative techniques in quantitative and human perception benchmarks. Additionally, ZeroComp extends to real and outdoor image compositing, even when trained solely on synthetic indoor data, showcasing its effectiveness in image compositing.
title ZeroComp: Zero-shot Object Compositing from Image Intrinsics via Diffusion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.08168