T2Bs: Text-to-Character Blendshapes via Video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Luo, Jiahao, Wang, Chaoyang, Vasilkovsky, Michael, Shakhrai, Vladislav, Liu, Di, Zhuang, Peiye, Tulyakov, Sergey, Wonka, Peter, Lee, Hsin-Ying, Davis, James, Wang, Jian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916972768264192
author Luo, Jiahao
Wang, Chaoyang
Vasilkovsky, Michael
Shakhrai, Vladislav
Liu, Di
Zhuang, Peiye
Tulyakov, Sergey
Wonka, Peter
Lee, Hsin-Ying
Davis, James
Wang, Jian
author_facet Luo, Jiahao
Wang, Chaoyang
Vasilkovsky, Michael
Shakhrai, Vladislav
Liu, Di
Zhuang, Peiye
Tulyakov, Sergey
Wonka, Peter
Lee, Hsin-Ying
Davis, James
Wang, Jian
contents We present T2Bs, a framework for generating high-quality, animatable character head morphable models from text by combining static text-to-3D generation with video diffusion. Text-to-3D models produce detailed static geometry but lack motion synthesis, while video diffusion models generate motion with temporal and multi-view geometric inconsistencies. T2Bs bridges this gap by leveraging deformable 3D Gaussian splatting to align static 3D assets with video outputs. By constraining motion with static geometry and employing a view-dependent deformation MLP, T2Bs (i) outperforms existing 4D generation methods in accuracy and expressiveness while reducing video artifacts and view inconsistencies, and (ii) reconstructs smooth, coherent, fully registered 3D geometries designed to scale for building morphable models with diverse, realistic facial motions. This enables synthesizing expressive, animatable character heads that surpass current 4D generation techniques.
format Preprint
id arxiv_https___arxiv_org_abs_2509_10678
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle T2Bs: Text-to-Character Blendshapes via Video Generation
Luo, Jiahao
Wang, Chaoyang
Vasilkovsky, Michael
Shakhrai, Vladislav
Liu, Di
Zhuang, Peiye
Tulyakov, Sergey
Wonka, Peter
Lee, Hsin-Ying
Davis, James
Wang, Jian
Graphics
We present T2Bs, a framework for generating high-quality, animatable character head morphable models from text by combining static text-to-3D generation with video diffusion. Text-to-3D models produce detailed static geometry but lack motion synthesis, while video diffusion models generate motion with temporal and multi-view geometric inconsistencies. T2Bs bridges this gap by leveraging deformable 3D Gaussian splatting to align static 3D assets with video outputs. By constraining motion with static geometry and employing a view-dependent deformation MLP, T2Bs (i) outperforms existing 4D generation methods in accuracy and expressiveness while reducing video artifacts and view inconsistencies, and (ii) reconstructs smooth, coherent, fully registered 3D geometries designed to scale for building morphable models with diverse, realistic facial motions. This enables synthesizing expressive, animatable character heads that surpass current 4D generation techniques.
title T2Bs: Text-to-Character Blendshapes via Video Generation
topic Graphics
url https://arxiv.org/abs/2509.10678