KeySync: A Robust Approach for Leakage-free Lip Synchronization in High Resolution

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bigata, Antoni, Mira, Rodrigo, Bounareli, Stella, Stypułkowski, Michał, Vougioukas, Konstantinos, Petridis, Stavros, Pantic, Maja
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915268965433344
author Bigata, Antoni
Mira, Rodrigo
Bounareli, Stella
Stypułkowski, Michał
Vougioukas, Konstantinos
Petridis, Stavros
Pantic, Maja
author_facet Bigata, Antoni
Mira, Rodrigo
Bounareli, Stella
Stypułkowski, Michał
Vougioukas, Konstantinos
Petridis, Stavros
Pantic, Maja
contents Lip synchronization, known as the task of aligning lip movements in an existing video with new input audio, is typically framed as a simpler variant of audio-driven facial animation. However, as well as suffering from the usual issues in talking head generation (e.g., temporal consistency), lip synchronization presents significant new challenges such as expression leakage from the input video and facial occlusions, which can severely impact real-world applications like automated dubbing, but are often neglected in existing works. To address these shortcomings, we present KeySync, a two-stage framework that succeeds in solving the issue of temporal consistency, while also incorporating solutions for leakage and occlusions using a carefully designed masking strategy. We show that KeySync achieves state-of-the-art results in lip reconstruction and cross-synchronization, improving visual quality and reducing expression leakage according to LipLeak, our novel leakage metric. Furthermore, we demonstrate the effectiveness of our new masking approach in handling occlusions and validate our architectural choices through several ablation studies. Code and model weights can be found at https://antonibigata.github.io/KeySync.
format Preprint
id arxiv_https___arxiv_org_abs_2505_00497
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle KeySync: A Robust Approach for Leakage-free Lip Synchronization in High Resolution
Bigata, Antoni
Mira, Rodrigo
Bounareli, Stella
Stypułkowski, Michał
Vougioukas, Konstantinos
Petridis, Stavros
Pantic, Maja
Computer Vision and Pattern Recognition
Lip synchronization, known as the task of aligning lip movements in an existing video with new input audio, is typically framed as a simpler variant of audio-driven facial animation. However, as well as suffering from the usual issues in talking head generation (e.g., temporal consistency), lip synchronization presents significant new challenges such as expression leakage from the input video and facial occlusions, which can severely impact real-world applications like automated dubbing, but are often neglected in existing works. To address these shortcomings, we present KeySync, a two-stage framework that succeeds in solving the issue of temporal consistency, while also incorporating solutions for leakage and occlusions using a carefully designed masking strategy. We show that KeySync achieves state-of-the-art results in lip reconstruction and cross-synchronization, improving visual quality and reducing expression leakage according to LipLeak, our novel leakage metric. Furthermore, we demonstrate the effectiveness of our new masking approach in handling occlusions and validate our architectural choices through several ablation studies. Code and model weights can be found at https://antonibigata.github.io/KeySync.
title KeySync: A Robust Approach for Leakage-free Lip Synchronization in High Resolution
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.00497