SyncVoice: Towards Video Dubbing with Vision-Augmented Pretrained TTS Model

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wang, Kaidi, He, Yi, Guan, Wenhao, Wu, Weijie, Ding, Hongwu, Zhang, Xiong, Wu, Di, Meng, Meng, Luan, Jian, Li, Lin, Hong, Qingyang
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917125577244672
author Wang, Kaidi
He, Yi
Guan, Wenhao
Wu, Weijie
Ding, Hongwu
Zhang, Xiong
Wu, Di
Meng, Meng
Luan, Jian
Li, Lin
Hong, Qingyang
author_facet Wang, Kaidi
He, Yi
Guan, Wenhao
Wu, Weijie
Ding, Hongwu
Zhang, Xiong
Wu, Di
Meng, Meng
Luan, Jian
Li, Lin
Hong, Qingyang
contents Video dubbing aims to generate high-fidelity speech that is precisely temporally aligned with the visual content. Existing methods still suffer from limitations in speech naturalness and audio-visual synchronization, and are limited to monolingual settings. To address these challenges, we propose SyncVoice, a vision-augmented video dubbing framework built upon a pretrained text-to-speech (TTS) model. By fine-tuning the TTS model on audio-visual data, we achieve strong audiovisual consistency. We propose a Dual Speaker Encoder to effectively mitigate inter-language interference in cross-lingual speech synthesis and explore the application of video dubbing in video translation scenarios. Experimental results show that SyncVoice achieves high-fidelity speech generation with strong synchronization performance, demonstrating its potential in video dubbing tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2512_05126
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SyncVoice: Towards Video Dubbing with Vision-Augmented Pretrained TTS Model
Wang, Kaidi
He, Yi
Guan, Wenhao
Wu, Weijie
Ding, Hongwu
Zhang, Xiong
Wu, Di
Meng, Meng
Luan, Jian
Li, Lin
Hong, Qingyang
Audio and Speech Processing
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Multimedia
Sound
Video dubbing aims to generate high-fidelity speech that is precisely temporally aligned with the visual content. Existing methods still suffer from limitations in speech naturalness and audio-visual synchronization, and are limited to monolingual settings. To address these challenges, we propose SyncVoice, a vision-augmented video dubbing framework built upon a pretrained text-to-speech (TTS) model. By fine-tuning the TTS model on audio-visual data, we achieve strong audiovisual consistency. We propose a Dual Speaker Encoder to effectively mitigate inter-language interference in cross-lingual speech synthesis and explore the application of video dubbing in video translation scenarios. Experimental results show that SyncVoice achieves high-fidelity speech generation with strong synchronization performance, demonstrating its potential in video dubbing tasks.
title SyncVoice: Towards Video Dubbing with Vision-Augmented Pretrained TTS Model
topic Audio and Speech Processing
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Multimedia
Sound
url https://arxiv.org/abs/2512.05126