DAGSM: Disentangled Avatar Generation with GS-enhanced Mesh

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhuang, Jingyu, Kang, Di, Bao, Linchao, Lin, Liang, Li, Guanbin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909540329455616
author Zhuang, Jingyu
Kang, Di
Bao, Linchao
Lin, Liang
Li, Guanbin
author_facet Zhuang, Jingyu
Kang, Di
Bao, Linchao
Lin, Liang
Li, Guanbin
contents Text-driven avatar generation has gained significant attention owing to its convenience. However, existing methods typically model the human body with all garments as a single 3D model, limiting its usability, such as clothing replacement, and reducing user control over the generation process. To overcome the limitations above, we propose DAGSM, a novel pipeline that generates disentangled human bodies and garments from the given text prompts. Specifically, we model each part (e.g., body, upper/lower clothes) of the clothed human as one GS-enhanced mesh (GSM), which is a traditional mesh attached with 2D Gaussians to better handle complicated textures (e.g., woolen, translucent clothes) and produce realistic cloth animations. During the generation, we first create the unclothed body, followed by a sequence of individual cloth generation based on the body, where we introduce a semantic-based algorithm to achieve better human-cloth and garment-garment separation. To improve texture quality, we propose a view-consistent texture refinement module, including a cross-view attention mechanism for texture style consistency and an incident-angle-weighted denoising (IAW-DE) strategy to update the appearance. Extensive experiments have demonstrated that DAGSM generates high-quality disentangled avatars, supports clothing replacement and realistic animation, and outperforms the baselines in visual quality.
format Preprint
id arxiv_https___arxiv_org_abs_2411_15205
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DAGSM: Disentangled Avatar Generation with GS-enhanced Mesh
Zhuang, Jingyu
Kang, Di
Bao, Linchao
Lin, Liang
Li, Guanbin
Computer Vision and Pattern Recognition
Graphics
Text-driven avatar generation has gained significant attention owing to its convenience. However, existing methods typically model the human body with all garments as a single 3D model, limiting its usability, such as clothing replacement, and reducing user control over the generation process. To overcome the limitations above, we propose DAGSM, a novel pipeline that generates disentangled human bodies and garments from the given text prompts. Specifically, we model each part (e.g., body, upper/lower clothes) of the clothed human as one GS-enhanced mesh (GSM), which is a traditional mesh attached with 2D Gaussians to better handle complicated textures (e.g., woolen, translucent clothes) and produce realistic cloth animations. During the generation, we first create the unclothed body, followed by a sequence of individual cloth generation based on the body, where we introduce a semantic-based algorithm to achieve better human-cloth and garment-garment separation. To improve texture quality, we propose a view-consistent texture refinement module, including a cross-view attention mechanism for texture style consistency and an incident-angle-weighted denoising (IAW-DE) strategy to update the appearance. Extensive experiments have demonstrated that DAGSM generates high-quality disentangled avatars, supports clothing replacement and realistic animation, and outperforms the baselines in visual quality.
title DAGSM: Disentangled Avatar Generation with GS-enhanced Mesh
topic Computer Vision and Pattern Recognition
Graphics
url https://arxiv.org/abs/2411.15205