SyncVoice: Towards Video Dubbing with Vision-Augmented Pretrained TTS Model
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866917125577244672 |
|---|---|
| author | Wang, Kaidi He, Yi Guan, Wenhao Wu, Weijie Ding, Hongwu Zhang, Xiong Wu, Di Meng, Meng Luan, Jian Li, Lin Hong, Qingyang |
| author_facet | Wang, Kaidi He, Yi Guan, Wenhao Wu, Weijie Ding, Hongwu Zhang, Xiong Wu, Di Meng, Meng Luan, Jian Li, Lin Hong, Qingyang |
| contents | Video dubbing aims to generate high-fidelity speech that is precisely temporally aligned with the visual content. Existing methods still suffer from limitations in speech naturalness and audio-visual synchronization, and are limited to monolingual settings. To address these challenges, we propose SyncVoice, a vision-augmented video dubbing framework built upon a pretrained text-to-speech (TTS) model. By fine-tuning the TTS model on audio-visual data, we achieve strong audiovisual consistency. We propose a Dual Speaker Encoder to effectively mitigate inter-language interference in cross-lingual speech synthesis and explore the application of video dubbing in video translation scenarios. Experimental results show that SyncVoice achieves high-fidelity speech generation with strong synchronization performance, demonstrating its potential in video dubbing tasks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_05126 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | SyncVoice: Towards Video Dubbing with Vision-Augmented Pretrained TTS Model Wang, Kaidi He, Yi Guan, Wenhao Wu, Weijie Ding, Hongwu Zhang, Xiong Wu, Di Meng, Meng Luan, Jian Li, Lin Hong, Qingyang Audio and Speech Processing Artificial Intelligence Computation and Language Computer Vision and Pattern Recognition Multimedia Sound Video dubbing aims to generate high-fidelity speech that is precisely temporally aligned with the visual content. Existing methods still suffer from limitations in speech naturalness and audio-visual synchronization, and are limited to monolingual settings. To address these challenges, we propose SyncVoice, a vision-augmented video dubbing framework built upon a pretrained text-to-speech (TTS) model. By fine-tuning the TTS model on audio-visual data, we achieve strong audiovisual consistency. We propose a Dual Speaker Encoder to effectively mitigate inter-language interference in cross-lingual speech synthesis and explore the application of video dubbing in video translation scenarios. Experimental results show that SyncVoice achieves high-fidelity speech generation with strong synchronization performance, demonstrating its potential in video dubbing tasks. |
| title | SyncVoice: Towards Video Dubbing with Vision-Augmented Pretrained TTS Model |
| topic | Audio and Speech Processing Artificial Intelligence Computation and Language Computer Vision and Pattern Recognition Multimedia Sound |
| url | https://arxiv.org/abs/2512.05126 |