Towards a Training Free Approach for 3D Scene Editing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Madhavaram, Vivek, Rawat, Shivangana, Devaguptapu, Chaitanya, Sharma, Charu, Kaul, Manohar
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915068929638400
author Madhavaram, Vivek
Rawat, Shivangana
Devaguptapu, Chaitanya
Sharma, Charu
Kaul, Manohar
author_facet Madhavaram, Vivek
Rawat, Shivangana
Devaguptapu, Chaitanya
Sharma, Charu
Kaul, Manohar
contents Text driven diffusion models have shown remarkable capabilities in editing images. However, when editing 3D scenes, existing works mostly rely on training a NeRF for 3D editing. Recent NeRF editing methods leverages edit operations by deploying 2D diffusion models and project these edits into 3D space. They require strong positional priors alongside text prompt to identify the edit location. These methods are operational on small 3D scenes and are more generalized to particular scene. They require training for each specific edit and cannot be exploited in real-time edits. To address these limitations, we propose a novel method, FreeEdit, to make edits in training free manner using mesh representations as a substitute for NeRF. Training-free methods are now a possibility because of the advances in foundation model's space. We leverage these models to bring a training-free alternative and introduce solutions for insertion, replacement and deletion. We consider insertion, replacement and deletion as basic blocks for performing intricate edits with certain combinations of these operations. Given a text prompt and a 3D scene, our model is capable of identifying what object should be inserted/replaced or deleted and location where edit should be performed. We also introduce a novel algorithm as part of FreeEdit to find the optimal location on grounding object for placement. We evaluate our model by comparing it with baseline models on a wide range of scenes using quantitative and qualitative metrics and showcase the merits of our method with respect to others.
format Preprint
id arxiv_https___arxiv_org_abs_2412_12766
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Towards a Training Free Approach for 3D Scene Editing
Madhavaram, Vivek
Rawat, Shivangana
Devaguptapu, Chaitanya
Sharma, Charu
Kaul, Manohar
Computer Vision and Pattern Recognition
Text driven diffusion models have shown remarkable capabilities in editing images. However, when editing 3D scenes, existing works mostly rely on training a NeRF for 3D editing. Recent NeRF editing methods leverages edit operations by deploying 2D diffusion models and project these edits into 3D space. They require strong positional priors alongside text prompt to identify the edit location. These methods are operational on small 3D scenes and are more generalized to particular scene. They require training for each specific edit and cannot be exploited in real-time edits. To address these limitations, we propose a novel method, FreeEdit, to make edits in training free manner using mesh representations as a substitute for NeRF. Training-free methods are now a possibility because of the advances in foundation model's space. We leverage these models to bring a training-free alternative and introduce solutions for insertion, replacement and deletion. We consider insertion, replacement and deletion as basic blocks for performing intricate edits with certain combinations of these operations. Given a text prompt and a 3D scene, our model is capable of identifying what object should be inserted/replaced or deleted and location where edit should be performed. We also introduce a novel algorithm as part of FreeEdit to find the optimal location on grounding object for placement. We evaluate our model by comparing it with baseline models on a wide range of scenes using quantitative and qualitative metrics and showcase the merits of our method with respect to others.
title Towards a Training Free Approach for 3D Scene Editing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.12766