COMOGen: A Controllable Text-to-3D Multi-object Generation Framework

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Sun, Shaorong, Pang, Shuchao, Yao, Yazhou, Huang, Xiaoshui
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914932005535744
author Sun, Shaorong
Pang, Shuchao
Yao, Yazhou
Huang, Xiaoshui
author_facet Sun, Shaorong
Pang, Shuchao
Yao, Yazhou
Huang, Xiaoshui
contents The controllability of 3D object generation methods is achieved through input text. Existing text-to-3D object generation methods primarily focus on generating a single object based on a single object description. However, these methods often face challenges in producing results that accurately correspond to our desired positions when the input text involves multiple objects. To address the issue of controllability in generating multiple objects, this paper introduces COMOGen, a COntrollable text-to-3D Multi-Object Generation framework. COMOGen enables the simultaneous generation of multiple 3D objects by the distillation of layout and multi-view prior knowledge. The framework consists of three modules: the layout control module, the multi-view consistency control module, and the 3D content enhancement module. Moreover, to integrate these three modules as an integral framework, we propose Layout Multi-view Score Distillation, which unifies two prior knowledge and further enhances the diversity and quality of generated 3D content. Comprehensive experiments demonstrate the effectiveness of our approach compared to the state-of-the-art methods, which represents a significant step forward in enabling more controlled and versatile text-based 3D content generation.
format Preprint
id arxiv_https___arxiv_org_abs_2409_00590
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle COMOGen: A Controllable Text-to-3D Multi-object Generation Framework
Sun, Shaorong
Pang, Shuchao
Yao, Yazhou
Huang, Xiaoshui
Computer Vision and Pattern Recognition
The controllability of 3D object generation methods is achieved through input text. Existing text-to-3D object generation methods primarily focus on generating a single object based on a single object description. However, these methods often face challenges in producing results that accurately correspond to our desired positions when the input text involves multiple objects. To address the issue of controllability in generating multiple objects, this paper introduces COMOGen, a COntrollable text-to-3D Multi-Object Generation framework. COMOGen enables the simultaneous generation of multiple 3D objects by the distillation of layout and multi-view prior knowledge. The framework consists of three modules: the layout control module, the multi-view consistency control module, and the 3D content enhancement module. Moreover, to integrate these three modules as an integral framework, we propose Layout Multi-view Score Distillation, which unifies two prior knowledge and further enhances the diversity and quality of generated 3D content. Comprehensive experiments demonstrate the effectiveness of our approach compared to the state-of-the-art methods, which represents a significant step forward in enabling more controlled and versatile text-based 3D content generation.
title COMOGen: A Controllable Text-to-3D Multi-object Generation Framework
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2409.00590