FlexAvatar: Flexible Large Reconstruction Model for Animatable Gaussian Head Avatars with Detailed Deformation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Peng, Cheng, Su, Zhuo, Wang, Liao, Guo, Chen, Li, Zhaohu, Long, Chengjiang, Lv, Zheng, Sun, Jingxiang, Zhang, Chenyangguang, Liu, Yebin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914210649210880
author Peng, Cheng
Su, Zhuo
Wang, Liao
Guo, Chen
Li, Zhaohu
Long, Chengjiang
Lv, Zheng
Sun, Jingxiang
Zhang, Chenyangguang
Liu, Yebin
author_facet Peng, Cheng
Su, Zhuo
Wang, Liao
Guo, Chen
Li, Zhaohu
Long, Chengjiang
Lv, Zheng
Sun, Jingxiang
Zhang, Chenyangguang
Liu, Yebin
contents We present FlexAvatar, a flexible large reconstruction model for high-fidelity 3D head avatars with detailed dynamic deformation from single or sparse images, without requiring camera poses or expression labels. It leverages a transformer-based reconstruction model with structured head query tokens as canonical anchor to aggregate flexible input-number-agnostic, camera-pose-free and expression-free inputs into a robust canonical 3D representation. For detailed dynamic deformation, we introduce a lightweight UNet decoder conditioned on UV-space position maps, which can produce detailed expression-dependent deformations in real time. To better capture rare but critical expressions like wrinkles and bared teeth, we also adopt a data distribution adjustment strategy during training to balance the distribution of these expressions in the training set. Moreover, a lightweight 10-second refinement can further enhances identity-specific details in extreme identities without affecting deformation quality. Extensive experiments demonstrate that our FlexAvatar achieves superior 3D consistency, detailed dynamic realism compared with previous methods, providing a practical solution for animatable 3D avatar creation.
format Preprint
id arxiv_https___arxiv_org_abs_2512_17717
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FlexAvatar: Flexible Large Reconstruction Model for Animatable Gaussian Head Avatars with Detailed Deformation
Peng, Cheng
Su, Zhuo
Wang, Liao
Guo, Chen
Li, Zhaohu
Long, Chengjiang
Lv, Zheng
Sun, Jingxiang
Zhang, Chenyangguang
Liu, Yebin
Computer Vision and Pattern Recognition
We present FlexAvatar, a flexible large reconstruction model for high-fidelity 3D head avatars with detailed dynamic deformation from single or sparse images, without requiring camera poses or expression labels. It leverages a transformer-based reconstruction model with structured head query tokens as canonical anchor to aggregate flexible input-number-agnostic, camera-pose-free and expression-free inputs into a robust canonical 3D representation. For detailed dynamic deformation, we introduce a lightweight UNet decoder conditioned on UV-space position maps, which can produce detailed expression-dependent deformations in real time. To better capture rare but critical expressions like wrinkles and bared teeth, we also adopt a data distribution adjustment strategy during training to balance the distribution of these expressions in the training set. Moreover, a lightweight 10-second refinement can further enhances identity-specific details in extreme identities without affecting deformation quality. Extensive experiments demonstrate that our FlexAvatar achieves superior 3D consistency, detailed dynamic realism compared with previous methods, providing a practical solution for animatable 3D avatar creation.
title FlexAvatar: Flexible Large Reconstruction Model for Animatable Gaussian Head Avatars with Detailed Deformation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.17717