VEIGAR: View-consistent Explicit Inpainting and Geometry Alignment for 3D object Removal

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Do, Pham Khai Nguyen, Tran, Bao Nguyen, Nguyen, Nam, Nguyen, Duc Dung
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911013258919936
author Do, Pham Khai Nguyen
Tran, Bao Nguyen
Nguyen, Nam
Nguyen, Duc Dung
author_facet Do, Pham Khai Nguyen
Tran, Bao Nguyen
Nguyen, Nam
Nguyen, Duc Dung
contents Recent advances in Novel View Synthesis (NVS) and 3D generation have significantly improved editing tasks, with a primary emphasis on maintaining cross-view consistency throughout the generative process. Contemporary methods typically address this challenge using a dual-strategy framework: performing consistent 2D inpainting across all views guided by embedded priors either explicitly in pixel space or implicitly in latent space; and conducting 3D reconstruction with additional consistency guidance. Previous strategies, in particular, often require an initial 3D reconstruction phase to establish geometric structure, introducing considerable computational overhead. Even with the added cost, the resulting reconstruction quality often remains suboptimal. In this paper, we present VEIGAR, a computationally efficient framework that outperforms existing methods without relying on an initial reconstruction phase. VEIGAR leverages a lightweight foundation model to reliably align priors explicitly in the pixel space. In addition, we introduce a novel supervision strategy based on scale-invariant depth loss, which removes the need for traditional scale-and-shift operations in monocular depth regularization. Through extensive experimentation, VEIGAR establishes a new state-of-the-art benchmark in reconstruction quality and cross-view consistency, while achieving a threefold reduction in training time compared to the fastest existing method, highlighting its superior balance of efficiency and effectiveness.
format Preprint
id arxiv_https___arxiv_org_abs_2506_15821
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VEIGAR: View-consistent Explicit Inpainting and Geometry Alignment for 3D object Removal
Do, Pham Khai Nguyen
Tran, Bao Nguyen
Nguyen, Nam
Nguyen, Duc Dung
Graphics
Artificial Intelligence
Computer Vision and Pattern Recognition
Image and Video Processing
Recent advances in Novel View Synthesis (NVS) and 3D generation have significantly improved editing tasks, with a primary emphasis on maintaining cross-view consistency throughout the generative process. Contemporary methods typically address this challenge using a dual-strategy framework: performing consistent 2D inpainting across all views guided by embedded priors either explicitly in pixel space or implicitly in latent space; and conducting 3D reconstruction with additional consistency guidance. Previous strategies, in particular, often require an initial 3D reconstruction phase to establish geometric structure, introducing considerable computational overhead. Even with the added cost, the resulting reconstruction quality often remains suboptimal. In this paper, we present VEIGAR, a computationally efficient framework that outperforms existing methods without relying on an initial reconstruction phase. VEIGAR leverages a lightweight foundation model to reliably align priors explicitly in the pixel space. In addition, we introduce a novel supervision strategy based on scale-invariant depth loss, which removes the need for traditional scale-and-shift operations in monocular depth regularization. Through extensive experimentation, VEIGAR establishes a new state-of-the-art benchmark in reconstruction quality and cross-view consistency, while achieving a threefold reduction in training time compared to the fastest existing method, highlighting its superior balance of efficiency and effectiveness.
title VEIGAR: View-consistent Explicit Inpainting and Geometry Alignment for 3D object Removal
topic Graphics
Artificial Intelligence
Computer Vision and Pattern Recognition
Image and Video Processing
url https://arxiv.org/abs/2506.15821