NeRF-Insert: 3D Local Editing with Multimodal Control Signals
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917654126657536 |
|---|---|
| author | Sabat, Benet Oriol Achille, Alessandro Trager, Matthew Soatto, Stefano |
| author_facet | Sabat, Benet Oriol Achille, Alessandro Trager, Matthew Soatto, Stefano |
| contents | We propose NeRF-Insert, a NeRF editing framework that allows users to make high-quality local edits with a flexible level of control. Unlike previous work that relied on image-to-image models, we cast scene editing as an in-painting problem, which encourages the global structure of the scene to be preserved. Moreover, while most existing methods use only textual prompts to condition edits, our framework accepts a combination of inputs of different modalities as reference. More precisely, a user may provide a combination of textual and visual inputs including images, CAD models, and binary image masks for specifying a 3D region. We use generic image generation models to in-paint the scene from multiple viewpoints, and lift the local edits to a 3D-consistent NeRF edit. Compared to previous methods, our results show better visual quality and also maintain stronger consistency with the original NeRF. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2404_19204 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | NeRF-Insert: 3D Local Editing with Multimodal Control Signals Sabat, Benet Oriol Achille, Alessandro Trager, Matthew Soatto, Stefano Computer Vision and Pattern Recognition Artificial Intelligence Graphics We propose NeRF-Insert, a NeRF editing framework that allows users to make high-quality local edits with a flexible level of control. Unlike previous work that relied on image-to-image models, we cast scene editing as an in-painting problem, which encourages the global structure of the scene to be preserved. Moreover, while most existing methods use only textual prompts to condition edits, our framework accepts a combination of inputs of different modalities as reference. More precisely, a user may provide a combination of textual and visual inputs including images, CAD models, and binary image masks for specifying a 3D region. We use generic image generation models to in-paint the scene from multiple viewpoints, and lift the local edits to a 3D-consistent NeRF edit. Compared to previous methods, our results show better visual quality and also maintain stronger consistency with the original NeRF. |
| title | NeRF-Insert: 3D Local Editing with Multimodal Control Signals |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Graphics |
| url | https://arxiv.org/abs/2404.19204 |