TeRA: Rethinking Text-guided Realistic 3D Avatar Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Yanwen, Zhuang, Yiyu, Zhang, Jiawei, Wang, Li, Zeng, Yifei, Cao, Xun, Zuo, Xinxin, Zhu, Hao
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915476900151296
author Wang, Yanwen
Zhuang, Yiyu
Zhang, Jiawei
Wang, Li
Zeng, Yifei
Cao, Xun
Zuo, Xinxin
Zhu, Hao
author_facet Wang, Yanwen
Zhuang, Yiyu
Zhang, Jiawei
Wang, Li
Zeng, Yifei
Cao, Xun
Zuo, Xinxin
Zhu, Hao
contents In this paper, we rethink text-to-avatar generative models by proposing TeRA, a more efficient and effective framework than the previous SDS-based models and general large 3D generative models. Our approach employs a two-stage training strategy for learning a native 3D avatar generative model. Initially, we distill a decoder to derive a structured latent space from a large human reconstruction model. Subsequently, a text-controlled latent diffusion model is trained to generate photorealistic 3D human avatars within this latent space. TeRA enhances the model performance by eliminating slow iterative optimization and enables text-based partial customization through a structured 3D human representation. Experiments have proven our approach's superiority over previous text-to-avatar generative models in subjective and objective evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2509_02466
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TeRA: Rethinking Text-guided Realistic 3D Avatar Generation
Wang, Yanwen
Zhuang, Yiyu
Zhang, Jiawei
Wang, Li
Zeng, Yifei
Cao, Xun
Zuo, Xinxin
Zhu, Hao
Computer Vision and Pattern Recognition
In this paper, we rethink text-to-avatar generative models by proposing TeRA, a more efficient and effective framework than the previous SDS-based models and general large 3D generative models. Our approach employs a two-stage training strategy for learning a native 3D avatar generative model. Initially, we distill a decoder to derive a structured latent space from a large human reconstruction model. Subsequently, a text-controlled latent diffusion model is trained to generate photorealistic 3D human avatars within this latent space. TeRA enhances the model performance by eliminating slow iterative optimization and enables text-based partial customization through a structured 3D human representation. Experiments have proven our approach's superiority over previous text-to-avatar generative models in subjective and objective evaluation.
title TeRA: Rethinking Text-guided Realistic 3D Avatar Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.02466