StdGEN++: A Comprehensive System for Semantic-Decomposed 3D Character Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: He, Yuze, Zhou, Yanning, Zhao, Wang, Ye, Jingwen, Wu, Zhongkai, Yi, Ran, Liu, Yong-Jin
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918284079661056
author He, Yuze
Zhou, Yanning
Zhao, Wang
Ye, Jingwen
Wu, Zhongkai
Yi, Ran
Liu, Yong-Jin
author_facet He, Yuze
Zhou, Yanning
Zhao, Wang
Ye, Jingwen
Wu, Zhongkai
Yi, Ran
Liu, Yong-Jin
contents We present StdGEN++, a novel and comprehensive system for generating high-fidelity, semantically decomposed 3D characters from diverse inputs. Existing 3D generative methods often produce monolithic meshes that lack the structural flexibility required by industrial pipelines in gaming and animation. Addressing this gap, StdGEN++ is built upon a Dual-branch Semantic-aware Large Reconstruction Model (Dual-Branch S-LRM), which jointly reconstructs geometry, color, and per-component semantics in a feed-forward manner. To achieve production-level fidelity, we introduce a novel semantic surface extraction formalism compatible with hybrid implicit fields. This mechanism is accelerated by a coarse-to-fine proposal scheme, which significantly reduces memory footprint and enables high-resolution mesh generation. Furthermore, we propose a video-diffusion-based texture decomposition module that disentangles appearance into editable layers (e.g., separated iris and skin), resolving semantic confusion in facial regions. Experiments demonstrate that StdGEN++ achieves state-of-the-art performance, significantly outperforming existing methods in geometric accuracy and semantic disentanglement. Crucially, the resulting structural independence unlocks advanced downstream capabilities, including non-destructive editing, physics-compliant animation, and gaze tracking, making it a robust solution for automated character asset production.
format Preprint
id arxiv_https___arxiv_org_abs_2601_07660
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle StdGEN++: A Comprehensive System for Semantic-Decomposed 3D Character Generation
He, Yuze
Zhou, Yanning
Zhao, Wang
Ye, Jingwen
Wu, Zhongkai
Yi, Ran
Liu, Yong-Jin
Computer Vision and Pattern Recognition
We present StdGEN++, a novel and comprehensive system for generating high-fidelity, semantically decomposed 3D characters from diverse inputs. Existing 3D generative methods often produce monolithic meshes that lack the structural flexibility required by industrial pipelines in gaming and animation. Addressing this gap, StdGEN++ is built upon a Dual-branch Semantic-aware Large Reconstruction Model (Dual-Branch S-LRM), which jointly reconstructs geometry, color, and per-component semantics in a feed-forward manner. To achieve production-level fidelity, we introduce a novel semantic surface extraction formalism compatible with hybrid implicit fields. This mechanism is accelerated by a coarse-to-fine proposal scheme, which significantly reduces memory footprint and enables high-resolution mesh generation. Furthermore, we propose a video-diffusion-based texture decomposition module that disentangles appearance into editable layers (e.g., separated iris and skin), resolving semantic confusion in facial regions. Experiments demonstrate that StdGEN++ achieves state-of-the-art performance, significantly outperforming existing methods in geometric accuracy and semantic disentanglement. Crucially, the resulting structural independence unlocks advanced downstream capabilities, including non-destructive editing, physics-compliant animation, and gaze tracking, making it a robust solution for automated character asset production.
title StdGEN++: A Comprehensive System for Semantic-Decomposed 3D Character Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.07660