Saved in:
Bibliographic Details
Main Authors: Tan, Weipeng, Lin, Chuming, Xu, Chengming, Ji, Xiaozhong, Zhu, Junwei, Wang, Chengjie, Wu, Yunsheng, Fu, Yanwei
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2409.03270
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929608178270208
author Tan, Weipeng
Lin, Chuming
Xu, Chengming
Ji, Xiaozhong
Zhu, Junwei
Wang, Chengjie
Wu, Yunsheng
Fu, Yanwei
author_facet Tan, Weipeng
Lin, Chuming
Xu, Chengming
Ji, Xiaozhong
Zhu, Junwei
Wang, Chengjie
Wu, Yunsheng
Fu, Yanwei
contents Talking Head Generation (THG), typically driven by audio, is an important and challenging task with broad application prospects in various fields such as digital humans, film production, and virtual reality. While diffusion model-based THG methods present high quality and stable content generation, they often overlook the intrinsic style which encompasses personalized features such as speaking habits and facial expressions of a video. As consequence, the generated video content lacks diversity and vividness, thus being limited in real life scenarios. To address these issues, we propose a novel framework named Style-Enhanced Vivid Portrait (SVP) which fully leverages style-related information in THG. Specifically, we first introduce the novel probabilistic style prior learning to model the intrinsic style as a Gaussian distribution using facial expressions and audio embedding. The distribution is learned through the 'bespoked' contrastive objective, effectively capturing the dynamic style information in each video. Then we finetune a pretrained Stable Diffusion (SD) model to inject the learned intrinsic style as a controlling signal via cross attention. Experiments show that our model generates diverse, vivid, and high-quality videos with flexible control over intrinsic styles, outperforming existing state-of-the-art methods.
format Preprint
id arxiv_https___arxiv_org_abs_2409_03270
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SVP: Style-Enhanced Vivid Portrait Talking Head Diffusion Model
Tan, Weipeng
Lin, Chuming
Xu, Chengming
Ji, Xiaozhong
Zhu, Junwei
Wang, Chengjie
Wu, Yunsheng
Fu, Yanwei
Computer Vision and Pattern Recognition
Talking Head Generation (THG), typically driven by audio, is an important and challenging task with broad application prospects in various fields such as digital humans, film production, and virtual reality. While diffusion model-based THG methods present high quality and stable content generation, they often overlook the intrinsic style which encompasses personalized features such as speaking habits and facial expressions of a video. As consequence, the generated video content lacks diversity and vividness, thus being limited in real life scenarios. To address these issues, we propose a novel framework named Style-Enhanced Vivid Portrait (SVP) which fully leverages style-related information in THG. Specifically, we first introduce the novel probabilistic style prior learning to model the intrinsic style as a Gaussian distribution using facial expressions and audio embedding. The distribution is learned through the 'bespoked' contrastive objective, effectively capturing the dynamic style information in each video. Then we finetune a pretrained Stable Diffusion (SD) model to inject the learned intrinsic style as a controlling signal via cross attention. Experiments show that our model generates diverse, vivid, and high-quality videos with flexible control over intrinsic styles, outperforming existing state-of-the-art methods.
title SVP: Style-Enhanced Vivid Portrait Talking Head Diffusion Model
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2409.03270