CMMD: Contrastive Multi-Modal Diffusion for Video-Audio Conditional Modeling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Ruihan, Gamper, Hannes, Braun, Sebastian
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916429411909632
author Yang, Ruihan
Gamper, Hannes
Braun, Sebastian
author_facet Yang, Ruihan
Gamper, Hannes
Braun, Sebastian
contents We introduce a multi-modal diffusion model tailored for the bi-directional conditional generation of video and audio. We propose a joint contrastive training loss to improve the synchronization between visual and auditory occurrences. We present experiments on two datasets to evaluate the efficacy of our proposed model. The assessment of generation quality and alignment performance is carried out from various angles, encompassing both objective and subjective metrics. Our findings demonstrate that the proposed model outperforms the baseline in terms of quality and generation speed through introduction of our novel cross-modal easy fusion architectural block. Furthermore, the incorporation of the contrastive loss results in improvements in audio-visual alignment, particularly in the high-correlation video-to-audio generation task.
format Preprint
id arxiv_https___arxiv_org_abs_2312_05412
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle CMMD: Contrastive Multi-Modal Diffusion for Video-Audio Conditional Modeling
Yang, Ruihan
Gamper, Hannes
Braun, Sebastian
Machine Learning
Computer Vision and Pattern Recognition
Multimedia
Sound
Audio and Speech Processing
We introduce a multi-modal diffusion model tailored for the bi-directional conditional generation of video and audio. We propose a joint contrastive training loss to improve the synchronization between visual and auditory occurrences. We present experiments on two datasets to evaluate the efficacy of our proposed model. The assessment of generation quality and alignment performance is carried out from various angles, encompassing both objective and subjective metrics. Our findings demonstrate that the proposed model outperforms the baseline in terms of quality and generation speed through introduction of our novel cross-modal easy fusion architectural block. Furthermore, the incorporation of the contrastive loss results in improvements in audio-visual alignment, particularly in the high-correlation video-to-audio generation task.
title CMMD: Contrastive Multi-Modal Diffusion for Video-Audio Conditional Modeling
topic Machine Learning
Computer Vision and Pattern Recognition
Multimedia
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2312.05412