WordRobe: Text-Guided Generation of Textured 3D Garments

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Srivastava, Astitva, Manu, Pranav, Raj, Amit, Jampani, Varun, Sharma, Avinash
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909253002854400
author Srivastava, Astitva
Manu, Pranav
Raj, Amit
Jampani, Varun
Sharma, Avinash
author_facet Srivastava, Astitva
Manu, Pranav
Raj, Amit
Jampani, Varun
Sharma, Avinash
contents In this paper, we tackle a new and challenging problem of text-driven generation of 3D garments with high-quality textures. We propose "WordRobe", a novel framework for the generation of unposed & textured 3D garment meshes from user-friendly text prompts. We achieve this by first learning a latent representation of 3D garments using a novel coarse-to-fine training strategy and a loss for latent disentanglement, promoting better latent interpolation. Subsequently, we align the garment latent space to the CLIP embedding space in a weakly supervised manner, enabling text-driven 3D garment generation and editing. For appearance modeling, we leverage the zero-shot generation capability of ControlNet to synthesize view-consistent texture maps in a single feed-forward inference step, thereby drastically decreasing the generation time as compared to existing methods. We demonstrate superior performance over current SOTAs for learning 3D garment latent space, garment interpolation, and text-driven texture synthesis, supported by quantitative evaluation and qualitative user study. The unposed 3D garment meshes generated using WordRobe can be directly fed to standard cloth simulation & animation pipelines without any post-processing.
format Preprint
id arxiv_https___arxiv_org_abs_2403_17541
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle WordRobe: Text-Guided Generation of Textured 3D Garments
Srivastava, Astitva
Manu, Pranav
Raj, Amit
Jampani, Varun
Sharma, Avinash
Computer Vision and Pattern Recognition
Graphics
In this paper, we tackle a new and challenging problem of text-driven generation of 3D garments with high-quality textures. We propose "WordRobe", a novel framework for the generation of unposed & textured 3D garment meshes from user-friendly text prompts. We achieve this by first learning a latent representation of 3D garments using a novel coarse-to-fine training strategy and a loss for latent disentanglement, promoting better latent interpolation. Subsequently, we align the garment latent space to the CLIP embedding space in a weakly supervised manner, enabling text-driven 3D garment generation and editing. For appearance modeling, we leverage the zero-shot generation capability of ControlNet to synthesize view-consistent texture maps in a single feed-forward inference step, thereby drastically decreasing the generation time as compared to existing methods. We demonstrate superior performance over current SOTAs for learning 3D garment latent space, garment interpolation, and text-driven texture synthesis, supported by quantitative evaluation and qualitative user study. The unposed 3D garment meshes generated using WordRobe can be directly fed to standard cloth simulation & animation pipelines without any post-processing.
title WordRobe: Text-Guided Generation of Textured 3D Garments
topic Computer Vision and Pattern Recognition
Graphics
url https://arxiv.org/abs/2403.17541