FT2TF: First-Person Statement Text-To-Talking Face Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Diao, Xingjian, Cheng, Ming, Barrios, Wayner, Jin, SouYoung
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913581347373056
author Diao, Xingjian
Cheng, Ming
Barrios, Wayner
Jin, SouYoung
author_facet Diao, Xingjian
Cheng, Ming
Barrios, Wayner
Jin, SouYoung
contents Talking face generation has gained immense popularity in the computer vision community, with various applications including AR, VR, teleconferencing, digital assistants, and avatars. Traditional methods are mainly audio-driven, which have to deal with the inevitable resource-intensive nature of audio storage and processing. To address such a challenge, we propose FT2TF - First-Person Statement Text-To-Talking Face Generation, a novel one-stage end-to-end pipeline for talking face generation driven by first-person statement text. Different from previous work, our model only leverages visual and textual information without any other sources (e.g., audio/landmark/pose) during inference. Extensive experiments are conducted on LRS2 and LRS3 datasets, and results on multi-dimensional evaluation metrics are reported. Both quantitative and qualitative results showcase that FT2TF outperforms existing relevant methods and reaches the state-of-the-art. This achievement highlights our model's capability to bridge first-person statements and dynamic face generation, providing insightful guidance for future work.
format Preprint
id arxiv_https___arxiv_org_abs_2312_05430
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle FT2TF: First-Person Statement Text-To-Talking Face Generation
Diao, Xingjian
Cheng, Ming
Barrios, Wayner
Jin, SouYoung
Computer Vision and Pattern Recognition
Talking face generation has gained immense popularity in the computer vision community, with various applications including AR, VR, teleconferencing, digital assistants, and avatars. Traditional methods are mainly audio-driven, which have to deal with the inevitable resource-intensive nature of audio storage and processing. To address such a challenge, we propose FT2TF - First-Person Statement Text-To-Talking Face Generation, a novel one-stage end-to-end pipeline for talking face generation driven by first-person statement text. Different from previous work, our model only leverages visual and textual information without any other sources (e.g., audio/landmark/pose) during inference. Extensive experiments are conducted on LRS2 and LRS3 datasets, and results on multi-dimensional evaluation metrics are reported. Both quantitative and qualitative results showcase that FT2TF outperforms existing relevant methods and reaches the state-of-the-art. This achievement highlights our model's capability to bridge first-person statements and dynamic face generation, providing insightful guidance for future work.
title FT2TF: First-Person Statement Text-To-Talking Face Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2312.05430