Staff View: :: Library Catalog

Saved in:

Bibliographic Details
Main Authors:	Junli, Deng, Yihao, Luo, Xueting, Yang, Siyou, Li, Wei, Wang, Jinyang, Guo, Ping, Shi
Format:	Preprint
Published:	2024
Subjects:	Computer Vision and Pattern Recognition
Online Access:	https://arxiv.org/abs/2409.09326
Tags:	Add Tag No Tags, Be the first to tag this record!

_version_	1866910603744903168
author	Junli, Deng Yihao, Luo Xueting, Yang Siyou, Li Wei, Wang Jinyang, Guo Ping, Shi
author_facet	Junli, Deng Yihao, Luo Xueting, Yang Siyou, Li Wei, Wang Jinyang, Guo Ping, Shi
contents	In the domain of photorealistic avatar generation, the fidelity of audio-driven lip motion synthesis is essential for realistic virtual interactions. Existing methods face two key challenges: a lack of vivacity due to limited diversity in generated lip poses and noticeable anamorphose motions caused by poor temporal coherence. To address these issues, we propose LawDNet, a novel deep-learning architecture enhancing lip synthesis through a Local Affine Warping Deformation mechanism. This mechanism models the intricate lip movements in response to the audio input by controllable non-linear warping fields. These fields consist of local affine transformations focused on abstract keypoints within deep feature maps, offering a novel universal paradigm for feature warping in networks. Additionally, LawDNet incorporates a dual-stream discriminator for improved frame-to-frame continuity and employs face normalization techniques to handle pose and scene variations. Extensive evaluations demonstrate LawDNet's superior robustness and lip movement dynamism performance compared to previous methods. The advancements presented in this paper, including the methodologies, training data, source codes, and pre-trained models, will be made accessible to the research community.
format	Preprint
id	arxiv_https___arxiv_org_abs_2409_09326
institution	arXiv
publishDate	2024
record_format	arxiv
spellingShingle	LawDNet: Enhanced Audio-Driven Lip Synthesis via Local Affine Warping Deformation Junli, Deng Yihao, Luo Xueting, Yang Siyou, Li Wei, Wang Jinyang, Guo Ping, Shi Computer Vision and Pattern Recognition In the domain of photorealistic avatar generation, the fidelity of audio-driven lip motion synthesis is essential for realistic virtual interactions. Existing methods face two key challenges: a lack of vivacity due to limited diversity in generated lip poses and noticeable anamorphose motions caused by poor temporal coherence. To address these issues, we propose LawDNet, a novel deep-learning architecture enhancing lip synthesis through a Local Affine Warping Deformation mechanism. This mechanism models the intricate lip movements in response to the audio input by controllable non-linear warping fields. These fields consist of local affine transformations focused on abstract keypoints within deep feature maps, offering a novel universal paradigm for feature warping in networks. Additionally, LawDNet incorporates a dual-stream discriminator for improved frame-to-frame continuity and employs face normalization techniques to handle pose and scene variations. Extensive evaluations demonstrate LawDNet's superior robustness and lip movement dynamism performance compared to previous methods. The advancements presented in this paper, including the methodologies, training data, source codes, and pre-trained models, will be made accessible to the research community.
title	LawDNet: Enhanced Audio-Driven Lip Synthesis via Local Affine Warping Deformation
topic	Computer Vision and Pattern Recognition
url	https://arxiv.org/abs/2409.09326

Similar Items