PersonaTalk: Bring Attention to Your Persona in Visual Dubbing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Longhao, Liang, Shuang, Ge, Zhipeng, Hu, Tianshu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913494003089408
author Zhang, Longhao
Liang, Shuang
Ge, Zhipeng
Hu, Tianshu
author_facet Zhang, Longhao
Liang, Shuang
Ge, Zhipeng
Hu, Tianshu
contents For audio-driven visual dubbing, it remains a considerable challenge to uphold and highlight speaker's persona while synthesizing accurate lip synchronization. Existing methods fall short of capturing speaker's unique speaking style or preserving facial details. In this paper, we present PersonaTalk, an attention-based two-stage framework, including geometry construction and face rendering, for high-fidelity and personalized visual dubbing. In the first stage, we propose a style-aware audio encoding module that injects speaking style into audio features through a cross-attention layer. The stylized audio features are then used to drive speaker's template geometry to obtain lip-synced geometries. In the second stage, a dual-attention face renderer is introduced to render textures for the target geometries. It consists of two parallel cross-attention layers, namely Lip-Attention and Face-Attention, which respectively sample textures from different reference frames to render the entire face. With our innovative design, intricate facial details can be well preserved. Comprehensive experiments and user studies demonstrate our advantages over other state-of-the-art methods in terms of visual quality, lip-sync accuracy and persona preservation. Furthermore, as a person-generic framework, PersonaTalk can achieve competitive performance as state-of-the-art person-specific methods. Project Page: https://grisoon.github.io/PersonaTalk/.
format Preprint
id arxiv_https___arxiv_org_abs_2409_05379
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle PersonaTalk: Bring Attention to Your Persona in Visual Dubbing
Zhang, Longhao
Liang, Shuang
Ge, Zhipeng
Hu, Tianshu
Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
For audio-driven visual dubbing, it remains a considerable challenge to uphold and highlight speaker's persona while synthesizing accurate lip synchronization. Existing methods fall short of capturing speaker's unique speaking style or preserving facial details. In this paper, we present PersonaTalk, an attention-based two-stage framework, including geometry construction and face rendering, for high-fidelity and personalized visual dubbing. In the first stage, we propose a style-aware audio encoding module that injects speaking style into audio features through a cross-attention layer. The stylized audio features are then used to drive speaker's template geometry to obtain lip-synced geometries. In the second stage, a dual-attention face renderer is introduced to render textures for the target geometries. It consists of two parallel cross-attention layers, namely Lip-Attention and Face-Attention, which respectively sample textures from different reference frames to render the entire face. With our innovative design, intricate facial details can be well preserved. Comprehensive experiments and user studies demonstrate our advantages over other state-of-the-art methods in terms of visual quality, lip-sync accuracy and persona preservation. Furthermore, as a person-generic framework, PersonaTalk can achieve competitive performance as state-of-the-art person-specific methods. Project Page: https://grisoon.github.io/PersonaTalk/.
title PersonaTalk: Bring Attention to Your Persona in Visual Dubbing
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
url https://arxiv.org/abs/2409.05379