TextToon: Real-Time Text Toonify Head Avatar from Single Video

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Song, Luchuan, Chen, Lele, Liu, Celong, Liu, Pinxin, Xu, Chenliang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909343334531072
author Song, Luchuan
Chen, Lele
Liu, Celong
Liu, Pinxin
Xu, Chenliang
author_facet Song, Luchuan
Chen, Lele
Liu, Celong
Liu, Pinxin
Xu, Chenliang
contents We propose TextToon, a method to generate a drivable toonified avatar. Given a short monocular video sequence and a written instruction about the avatar style, our model can generate a high-fidelity toonified avatar that can be driven in real-time by another video with arbitrary identities. Existing related works heavily rely on multi-view modeling to recover geometry via texture embeddings, presented in a static manner, leading to control limitations. The multi-view video input also makes it difficult to deploy these models in real-world applications. To address these issues, we adopt a conditional embedding Tri-plane to learn realistic and stylized facial representations in a Gaussian deformation field. Additionally, we expand the stylization capabilities of 3D Gaussian Splatting by introducing an adaptive pixel-translation neural network and leveraging patch-aware contrastive learning to achieve high-quality images. To push our work into consumer applications, we develop a real-time system that can operate at 48 FPS on a GPU machine and 15-18 FPS on a mobile machine. Extensive experiments demonstrate the efficacy of our approach in generating textual avatars over existing methods in terms of quality and real-time animation. Please refer to our project page for more details: https://songluchuan.github.io/TextToon/.
format Preprint
id arxiv_https___arxiv_org_abs_2410_07160
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle TextToon: Real-Time Text Toonify Head Avatar from Single Video
Song, Luchuan
Chen, Lele
Liu, Celong
Liu, Pinxin
Xu, Chenliang
Computer Vision and Pattern Recognition
Graphics
We propose TextToon, a method to generate a drivable toonified avatar. Given a short monocular video sequence and a written instruction about the avatar style, our model can generate a high-fidelity toonified avatar that can be driven in real-time by another video with arbitrary identities. Existing related works heavily rely on multi-view modeling to recover geometry via texture embeddings, presented in a static manner, leading to control limitations. The multi-view video input also makes it difficult to deploy these models in real-world applications. To address these issues, we adopt a conditional embedding Tri-plane to learn realistic and stylized facial representations in a Gaussian deformation field. Additionally, we expand the stylization capabilities of 3D Gaussian Splatting by introducing an adaptive pixel-translation neural network and leveraging patch-aware contrastive learning to achieve high-quality images. To push our work into consumer applications, we develop a real-time system that can operate at 48 FPS on a GPU machine and 15-18 FPS on a mobile machine. Extensive experiments demonstrate the efficacy of our approach in generating textual avatars over existing methods in terms of quality and real-time animation. Please refer to our project page for more details: https://songluchuan.github.io/TextToon/.
title TextToon: Real-Time Text Toonify Head Avatar from Single Video
topic Computer Vision and Pattern Recognition
Graphics
url https://arxiv.org/abs/2410.07160