3D Mesh Editing using Masked LRMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gao, Will, Wang, Dilin, Fan, Yuchen, Bozic, Aljaz, Stuyck, Tuur, Li, Zhengqin, Dong, Zhao, Ranjan, Rakesh, Sarafianos, Nikolaos
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909787485110272
author Gao, Will
Wang, Dilin
Fan, Yuchen
Bozic, Aljaz
Stuyck, Tuur
Li, Zhengqin
Dong, Zhao
Ranjan, Rakesh
Sarafianos, Nikolaos
author_facet Gao, Will
Wang, Dilin
Fan, Yuchen
Bozic, Aljaz
Stuyck, Tuur
Li, Zhengqin
Dong, Zhao
Ranjan, Rakesh
Sarafianos, Nikolaos
contents We present a novel approach to shape editing, building on recent progress in 3D reconstruction from multi-view images. We formulate shape editing as a conditional reconstruction problem, where the model must reconstruct the input shape with the exception of a specified 3D region, in which the geometry should be generated from the conditional signal. To this end, we train a conditional Large Reconstruction Model (LRM) for masked reconstruction, using multi-view consistent masks rendered from a randomly generated 3D occlusion, and using one clean viewpoint as the conditional signal. During inference, we manually define a 3D region to edit and provide an edited image from a canonical viewpoint to fill that region. We demonstrate that, in just a single forward pass, our method not only preserves the input geometry in the unmasked region through reconstruction capabilities on par with SoTA, but is also expressive enough to perform a variety of mesh edits from a single image guidance that past works struggle with, while being 2-10x faster than the top-performing prior work.
format Preprint
id arxiv_https___arxiv_org_abs_2412_08641
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle 3D Mesh Editing using Masked LRMs
Gao, Will
Wang, Dilin
Fan, Yuchen
Bozic, Aljaz
Stuyck, Tuur
Li, Zhengqin
Dong, Zhao
Ranjan, Rakesh
Sarafianos, Nikolaos
Computer Vision and Pattern Recognition
We present a novel approach to shape editing, building on recent progress in 3D reconstruction from multi-view images. We formulate shape editing as a conditional reconstruction problem, where the model must reconstruct the input shape with the exception of a specified 3D region, in which the geometry should be generated from the conditional signal. To this end, we train a conditional Large Reconstruction Model (LRM) for masked reconstruction, using multi-view consistent masks rendered from a randomly generated 3D occlusion, and using one clean viewpoint as the conditional signal. During inference, we manually define a 3D region to edit and provide an edited image from a canonical viewpoint to fill that region. We demonstrate that, in just a single forward pass, our method not only preserves the input geometry in the unmasked region through reconstruction capabilities on par with SoTA, but is also expressive enough to perform a variety of mesh edits from a single image guidance that past works struggle with, while being 2-10x faster than the top-performing prior work.
title 3D Mesh Editing using Masked LRMs
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.08641