SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Liu, Haohe, Lan, Gael Le, Mei, Xinhao, Ni, Zhaoheng, Kumar, Anurag, Nagaraja, Varun, Wang, Wenwu, Plumbley, Mark D., Shi, Yangyang, Chandra, Vikas
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913620424654848
author Liu, Haohe
Lan, Gael Le
Mei, Xinhao
Ni, Zhaoheng
Kumar, Anurag
Nagaraja, Varun
Wang, Wenwu
Plumbley, Mark D.
Shi, Yangyang
Chandra, Vikas
author_facet Liu, Haohe
Lan, Gael Le
Mei, Xinhao
Ni, Zhaoheng
Kumar, Anurag
Nagaraja, Varun
Wang, Wenwu
Plumbley, Mark D.
Shi, Yangyang
Chandra, Vikas
contents Video and audio are closely correlated modalities that humans naturally perceive together. While recent advancements have enabled the generation of audio or video from text, producing both modalities simultaneously still typically relies on either a cascaded process or multi-modal contrastive encoders. These approaches, however, often lead to suboptimal results due to inherent information losses during inference and conditioning. In this paper, we introduce SyncFlow, a system that is capable of simultaneously generating temporally synchronized audio and video from text. The core of SyncFlow is the proposed dual-diffusion-transformer (d-DiT) architecture, which enables joint video and audio modelling with proper information fusion. To efficiently manage the computational cost of joint audio and video modelling, SyncFlow utilizes a multi-stage training strategy that separates video and audio learning before joint fine-tuning. Our empirical evaluations demonstrate that SyncFlow produces audio and video outputs that are more correlated than baseline methods with significantly enhanced audio quality and audio-visual correspondence. Moreover, we demonstrate strong zero-shot capabilities of SyncFlow, including zero-shot video-to-audio generation and adaptation to novel video resolutions without further training.
format Preprint
id arxiv_https___arxiv_org_abs_2412_15220
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text
Liu, Haohe
Lan, Gael Le
Mei, Xinhao
Ni, Zhaoheng
Kumar, Anurag
Nagaraja, Varun
Wang, Wenwu
Plumbley, Mark D.
Shi, Yangyang
Chandra, Vikas
Multimedia
Sound
Audio and Speech Processing
Video and audio are closely correlated modalities that humans naturally perceive together. While recent advancements have enabled the generation of audio or video from text, producing both modalities simultaneously still typically relies on either a cascaded process or multi-modal contrastive encoders. These approaches, however, often lead to suboptimal results due to inherent information losses during inference and conditioning. In this paper, we introduce SyncFlow, a system that is capable of simultaneously generating temporally synchronized audio and video from text. The core of SyncFlow is the proposed dual-diffusion-transformer (d-DiT) architecture, which enables joint video and audio modelling with proper information fusion. To efficiently manage the computational cost of joint audio and video modelling, SyncFlow utilizes a multi-stage training strategy that separates video and audio learning before joint fine-tuning. Our empirical evaluations demonstrate that SyncFlow produces audio and video outputs that are more correlated than baseline methods with significantly enhanced audio quality and audio-visual correspondence. Moreover, we demonstrate strong zero-shot capabilities of SyncFlow, including zero-shot video-to-audio generation and adaptation to novel video resolutions without further training.
title SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text
topic Multimedia
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2412.15220