pix2gestalt: Amodal Segmentation by Synthesizing Wholes

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ozguroglu, Ege, Liu, Ruoshi, Surís, Dídac, Chen, Dian, Dave, Achal, Tokmakov, Pavel, Vondrick, Carl
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917574702268416
author Ozguroglu, Ege
Liu, Ruoshi
Surís, Dídac
Chen, Dian
Dave, Achal
Tokmakov, Pavel
Vondrick, Carl
author_facet Ozguroglu, Ege
Liu, Ruoshi
Surís, Dídac
Chen, Dian
Dave, Achal
Tokmakov, Pavel
Vondrick, Carl
contents We introduce pix2gestalt, a framework for zero-shot amodal segmentation, which learns to estimate the shape and appearance of whole objects that are only partially visible behind occlusions. By capitalizing on large-scale diffusion models and transferring their representations to this task, we learn a conditional diffusion model for reconstructing whole objects in challenging zero-shot cases, including examples that break natural and physical priors, such as art. As training data, we use a synthetically curated dataset containing occluded objects paired with their whole counterparts. Experiments show that our approach outperforms supervised baselines on established benchmarks. Our model can furthermore be used to significantly improve the performance of existing object recognition and 3D reconstruction methods in the presence of occlusions.
format Preprint
id arxiv_https___arxiv_org_abs_2401_14398
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle pix2gestalt: Amodal Segmentation by Synthesizing Wholes
Ozguroglu, Ege
Liu, Ruoshi
Surís, Dídac
Chen, Dian
Dave, Achal
Tokmakov, Pavel
Vondrick, Carl
Computer Vision and Pattern Recognition
Machine Learning
We introduce pix2gestalt, a framework for zero-shot amodal segmentation, which learns to estimate the shape and appearance of whole objects that are only partially visible behind occlusions. By capitalizing on large-scale diffusion models and transferring their representations to this task, we learn a conditional diffusion model for reconstructing whole objects in challenging zero-shot cases, including examples that break natural and physical priors, such as art. As training data, we use a synthetically curated dataset containing occluded objects paired with their whole counterparts. Experiments show that our approach outperforms supervised baselines on established benchmarks. Our model can furthermore be used to significantly improve the performance of existing object recognition and 3D reconstruction methods in the presence of occlusions.
title pix2gestalt: Amodal Segmentation by Synthesizing Wholes
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2401.14398