Saved in:
Bibliographic Details
Main Authors: Huang, Ding-Jiun, Wang, Yuanhao, Yuan, Shao-Ji, Mosella-Montoro, Albert, Carrasco, Francisco Vicente, Zhang, Cheng, De la Torre, Fernando
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2602.06122
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917252266196992
author Huang, Ding-Jiun
Wang, Yuanhao
Yuan, Shao-Ji
Mosella-Montoro, Albert
Carrasco, Francisco Vicente
Zhang, Cheng
De la Torre, Fernando
author_facet Huang, Ding-Jiun
Wang, Yuanhao
Yuan, Shao-Ji
Mosella-Montoro, Albert
Carrasco, Francisco Vicente
Zhang, Cheng
De la Torre, Fernando
contents Creating high-fidelity, animatable 3D talking heads is crucial for immersive applications, yet often hindered by the prevalence of low-quality image or video sources, which yield poor 3D reconstructions. In this paper, we introduce SuperHead, a novel framework for enhancing low-resolution, animatable 3D head avatars. The core challenge lies in synthesizing high-quality geometry and textures, while ensuring both 3D and temporal consistency during animation and preserving subject identity. Despite recent progress in image, video and 3D-based super-resolution (SR), existing SR techniques are ill-equipped to handle dynamic 3D inputs. To address this, SuperHead leverages the rich priors from pre-trained 3D generative models via a novel dynamics-aware 3D inversion scheme. This process optimizes the latent representation of the generative model to produce a super-resolved 3D Gaussian Splatting (3DGS) head model, which is subsequently rigged to an underlying parametric head model (e.g., FLAME) for animation. The inversion is jointly supervised using a sparse collection of upscaled 2D face renderings and corresponding depth maps, captured from diverse facial expressions and camera viewpoints, to ensure realism under dynamic facial motions. Experiments demonstrate that SuperHead generates avatars with fine-grained facial details under dynamic motions, significantly outperforming baseline methods in visual quality.
format Preprint
id arxiv_https___arxiv_org_abs_2602_06122
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle From Blurry to Believable: Enhancing Low-quality Talking Heads with 3D Generative Priors
Huang, Ding-Jiun
Wang, Yuanhao
Yuan, Shao-Ji
Mosella-Montoro, Albert
Carrasco, Francisco Vicente
Zhang, Cheng
De la Torre, Fernando
Computer Vision and Pattern Recognition
Creating high-fidelity, animatable 3D talking heads is crucial for immersive applications, yet often hindered by the prevalence of low-quality image or video sources, which yield poor 3D reconstructions. In this paper, we introduce SuperHead, a novel framework for enhancing low-resolution, animatable 3D head avatars. The core challenge lies in synthesizing high-quality geometry and textures, while ensuring both 3D and temporal consistency during animation and preserving subject identity. Despite recent progress in image, video and 3D-based super-resolution (SR), existing SR techniques are ill-equipped to handle dynamic 3D inputs. To address this, SuperHead leverages the rich priors from pre-trained 3D generative models via a novel dynamics-aware 3D inversion scheme. This process optimizes the latent representation of the generative model to produce a super-resolved 3D Gaussian Splatting (3DGS) head model, which is subsequently rigged to an underlying parametric head model (e.g., FLAME) for animation. The inversion is jointly supervised using a sparse collection of upscaled 2D face renderings and corresponding depth maps, captured from diverse facial expressions and camera viewpoints, to ensure realism under dynamic facial motions. Experiments demonstrate that SuperHead generates avatars with fine-grained facial details under dynamic motions, significantly outperforming baseline methods in visual quality.
title From Blurry to Believable: Enhancing Low-quality Talking Heads with 3D Generative Priors
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.06122