3DitScene: Editing Any Scene via Language-guided Disentangled Gaussian Splatting

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhang, Qihang, Xu, Yinghao, Wang, Chaoyang, Lee, Hsin-Ying, Wetzstein, Gordon, Zhou, Bolei, Yang, Ceyuan
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910461132275712
author Zhang, Qihang
Xu, Yinghao
Wang, Chaoyang
Lee, Hsin-Ying
Wetzstein, Gordon
Zhou, Bolei
Yang, Ceyuan
author_facet Zhang, Qihang
Xu, Yinghao
Wang, Chaoyang
Lee, Hsin-Ying
Wetzstein, Gordon
Zhou, Bolei
Yang, Ceyuan
contents Scene image editing is crucial for entertainment, photography, and advertising design. Existing methods solely focus on either 2D individual object or 3D global scene editing. This results in a lack of a unified approach to effectively control and manipulate scenes at the 3D level with different levels of granularity. In this work, we propose 3DitScene, a novel and unified scene editing framework leveraging language-guided disentangled Gaussian Splatting that enables seamless editing from 2D to 3D, allowing precise control over scene composition and individual objects. We first incorporate 3D Gaussians that are refined through generative priors and optimization techniques. Language features from CLIP then introduce semantics into 3D geometry for object disentanglement. With the disentangled Gaussians, 3DitScene allows for manipulation at both the global and individual levels, revolutionizing creative expression and empowering control over scenes and objects. Experimental results demonstrate the effectiveness and versatility of 3DitScene in scene image editing. Code and online demo can be found at our project homepage: https://zqh0253.github.io/3DitScene/.
format Preprint
id arxiv_https___arxiv_org_abs_2405_18424
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle 3DitScene: Editing Any Scene via Language-guided Disentangled Gaussian Splatting
Zhang, Qihang
Xu, Yinghao
Wang, Chaoyang
Lee, Hsin-Ying
Wetzstein, Gordon
Zhou, Bolei
Yang, Ceyuan
Computer Vision and Pattern Recognition
Scene image editing is crucial for entertainment, photography, and advertising design. Existing methods solely focus on either 2D individual object or 3D global scene editing. This results in a lack of a unified approach to effectively control and manipulate scenes at the 3D level with different levels of granularity. In this work, we propose 3DitScene, a novel and unified scene editing framework leveraging language-guided disentangled Gaussian Splatting that enables seamless editing from 2D to 3D, allowing precise control over scene composition and individual objects. We first incorporate 3D Gaussians that are refined through generative priors and optimization techniques. Language features from CLIP then introduce semantics into 3D geometry for object disentanglement. With the disentangled Gaussians, 3DitScene allows for manipulation at both the global and individual levels, revolutionizing creative expression and empowering control over scenes and objects. Experimental results demonstrate the effectiveness and versatility of 3DitScene in scene image editing. Code and online demo can be found at our project homepage: https://zqh0253.github.io/3DitScene/.
title 3DitScene: Editing Any Scene via Language-guided Disentangled Gaussian Splatting
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.18424