TextField3D: Towards Enhancing Open-Vocabulary 3D Generation with Noisy Text Fields

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Tianyu, Zeng, Yihan, Dong, Bowen, Xu, Hang, Xu, Songcen, Lau, Rynson W. H., Zuo, Wangmeng
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910366560157696
author Huang, Tianyu
Zeng, Yihan
Dong, Bowen
Xu, Hang
Xu, Songcen
Lau, Rynson W. H.
Zuo, Wangmeng
author_facet Huang, Tianyu
Zeng, Yihan
Dong, Bowen
Xu, Hang
Xu, Songcen
Lau, Rynson W. H.
Zuo, Wangmeng
contents Recent works learn 3D representation explicitly under text-3D guidance. However, limited text-3D data restricts the vocabulary scale and text control of generations. Generators may easily fall into a stereotype concept for certain text prompts, thus losing open-vocabulary generation ability. To tackle this issue, we introduce a conditional 3D generative model, namely TextField3D. Specifically, rather than using the text prompts as input directly, we suggest to inject dynamic noise into the latent space of given text prompts, i.e., Noisy Text Fields (NTFs). In this way, limited 3D data can be mapped to the appropriate range of textual latent space that is expanded by NTFs. To this end, an NTFGen module is proposed to model general text latent code in noisy fields. Meanwhile, an NTFBind module is proposed to align view-invariant image latent code to noisy fields, further supporting image-conditional 3D generation. To guide the conditional generation in both geometry and texture, multi-modal discrimination is constructed with a text-3D discriminator and a text-2.5D discriminator. Compared to previous methods, TextField3D includes three merits: 1) large vocabulary, 2) text consistency, and 3) low latency. Extensive experiments demonstrate that our method achieves a potential open-vocabulary 3D generation capability.
format Preprint
id arxiv_https___arxiv_org_abs_2309_17175
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle TextField3D: Towards Enhancing Open-Vocabulary 3D Generation with Noisy Text Fields
Huang, Tianyu
Zeng, Yihan
Dong, Bowen
Xu, Hang
Xu, Songcen
Lau, Rynson W. H.
Zuo, Wangmeng
Computer Vision and Pattern Recognition
Recent works learn 3D representation explicitly under text-3D guidance. However, limited text-3D data restricts the vocabulary scale and text control of generations. Generators may easily fall into a stereotype concept for certain text prompts, thus losing open-vocabulary generation ability. To tackle this issue, we introduce a conditional 3D generative model, namely TextField3D. Specifically, rather than using the text prompts as input directly, we suggest to inject dynamic noise into the latent space of given text prompts, i.e., Noisy Text Fields (NTFs). In this way, limited 3D data can be mapped to the appropriate range of textual latent space that is expanded by NTFs. To this end, an NTFGen module is proposed to model general text latent code in noisy fields. Meanwhile, an NTFBind module is proposed to align view-invariant image latent code to noisy fields, further supporting image-conditional 3D generation. To guide the conditional generation in both geometry and texture, multi-modal discrimination is constructed with a text-3D discriminator and a text-2.5D discriminator. Compared to previous methods, TextField3D includes three merits: 1) large vocabulary, 2) text consistency, and 3) low latency. Extensive experiments demonstrate that our method achieves a potential open-vocabulary 3D generation capability.
title TextField3D: Towards Enhancing Open-Vocabulary 3D Generation with Noisy Text Fields
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2309.17175