JUST-DUB-IT: Video Dubbing via Joint Audio-Visual Diffusion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Anthony, Korem, Naomi Ken, Zeevi, Gal, Halperin, Tavi, Yosef, Matan Ben, Jelercic, Urska, Bibi, Ofir, Patashnik, Or, Cohen-Or, Daniel
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909032468447232
author Chen, Anthony
Korem, Naomi Ken
Zeevi, Gal
Halperin, Tavi
Yosef, Matan Ben
Jelercic, Urska
Bibi, Ofir
Patashnik, Or
Cohen-Or, Daniel
author_facet Chen, Anthony
Korem, Naomi Ken
Zeevi, Gal
Halperin, Tavi
Yosef, Matan Ben
Jelercic, Urska
Bibi, Ofir
Patashnik, Or
Cohen-Or, Daniel
contents Audio-Visual Foundation Models, which are pretrained to jointly generate sound and visual content, have recently shown an unprecedented ability to model multi-modal generation and editing, opening new opportunities for downstream tasks. Among these tasks, video dubbing could greatly benefit from such priors, yet most existing solutions still rely on complex, task-specific pipelines that struggle in real-world settings. In this work, we introduce a single-model approach that adapts a foundational audio-video diffusion model for video-to-video dubbing via a lightweight LoRA. The LoRA enables the model to condition on an input audio-video while jointly generating translated audio and synchronized facial motion. To train this LoRA, we leverage the generative model itself to synthesize paired multilingual videos of the same speaker. Specifically, we generate multilingual videos with language switches within a single clip, and then inpaint the face and audio in each half to match the language of the other half. By leveraging the rich generative prior of the audio-visual model, our approach preserves speaker identity and lip synchronization while remaining robust to complex motion and real-world dynamics. We demonstrate that our approach produces high-quality dubbed videos with improved visual fidelity, lip synchronization, and robustness compared to existing dubbing pipelines.
format Preprint
id arxiv_https___arxiv_org_abs_2601_22143
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle JUST-DUB-IT: Video Dubbing via Joint Audio-Visual Diffusion
Chen, Anthony
Korem, Naomi Ken
Zeevi, Gal
Halperin, Tavi
Yosef, Matan Ben
Jelercic, Urska
Bibi, Ofir
Patashnik, Or
Cohen-Or, Daniel
Graphics
Computer Vision and Pattern Recognition
Audio-Visual Foundation Models, which are pretrained to jointly generate sound and visual content, have recently shown an unprecedented ability to model multi-modal generation and editing, opening new opportunities for downstream tasks. Among these tasks, video dubbing could greatly benefit from such priors, yet most existing solutions still rely on complex, task-specific pipelines that struggle in real-world settings. In this work, we introduce a single-model approach that adapts a foundational audio-video diffusion model for video-to-video dubbing via a lightweight LoRA. The LoRA enables the model to condition on an input audio-video while jointly generating translated audio and synchronized facial motion. To train this LoRA, we leverage the generative model itself to synthesize paired multilingual videos of the same speaker. Specifically, we generate multilingual videos with language switches within a single clip, and then inpaint the face and audio in each half to match the language of the other half. By leveraging the rich generative prior of the audio-visual model, our approach preserves speaker identity and lip synchronization while remaining robust to complex motion and real-world dynamics. We demonstrate that our approach produces high-quality dubbed videos with improved visual fidelity, lip synchronization, and robustness compared to existing dubbing pipelines.
title JUST-DUB-IT: Video Dubbing via Joint Audio-Visual Diffusion
topic Graphics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.22143