HighSync: High-Quality Lip Synchronization via Latent Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Daghigh, Saeed Firouzi, Mobarekeh, Majid Iranpour, Alavi, Mostafa, Bagheri, Mehdi
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914572315656192
author Daghigh, Saeed Firouzi
Mobarekeh, Majid Iranpour
Alavi, Mostafa
Bagheri, Mehdi
author_facet Daghigh, Saeed Firouzi
Mobarekeh, Majid Iranpour
Alavi, Mostafa
Bagheri, Mehdi
contents We present HighSync, an end-to-end diffusion-based framework for high-fidelity lip synchronization that generates photorealistic talking-face videos aligned with arbitrary input audio. Existing approaches consistently struggle to reconcile image quality with synchronization accuracy, producing either visually degraded outputs or temporally inconsistent lip movements. HighSync addresses both challenges simultaneously and, to our knowledge, is the first lip sync model to operate natively at 512*512 resolution, positioning it as a viable solution for professional production environments such as the film and broadcast industries. Central to our approach is the identification and systematic elimination of a data leakage phenomenon that has silently undermined temporal modeling in prior work, preventing models from developing a genuine dependence on the audio signal. Comprehensive evaluations across both perceptual quality and synchronization accuracy metrics confirm that HighSync achieves state-of-the-art performance on both fronts. Source code, pre-trained models, and supplementary video results are publicly available at: https://github.com/saeed5959/high_sync
format Preprint
id arxiv_https___arxiv_org_abs_2605_16918
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle HighSync: High-Quality Lip Synchronization via Latent Diffusion Models
Daghigh, Saeed Firouzi
Mobarekeh, Majid Iranpour
Alavi, Mostafa
Bagheri, Mehdi
Computer Vision and Pattern Recognition
We present HighSync, an end-to-end diffusion-based framework for high-fidelity lip synchronization that generates photorealistic talking-face videos aligned with arbitrary input audio. Existing approaches consistently struggle to reconcile image quality with synchronization accuracy, producing either visually degraded outputs or temporally inconsistent lip movements. HighSync addresses both challenges simultaneously and, to our knowledge, is the first lip sync model to operate natively at 512*512 resolution, positioning it as a viable solution for professional production environments such as the film and broadcast industries. Central to our approach is the identification and systematic elimination of a data leakage phenomenon that has silently undermined temporal modeling in prior work, preventing models from developing a genuine dependence on the audio signal. Comprehensive evaluations across both perceptual quality and synchronization accuracy metrics confirm that HighSync achieves state-of-the-art performance on both fronts. Source code, pre-trained models, and supplementary video results are publicly available at: https://github.com/saeed5959/high_sync
title HighSync: High-Quality Lip Synchronization via Latent Diffusion Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.16918