AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Haji-Ali, Moayed, Menapace, Willi, Siarohin, Aliaksandr, Skorokhodov, Ivan, Canberk, Alper, Lee, Kwot Sin, Ordonez, Vicente, Tulyakov, Sergey
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909532861497344
author Haji-Ali, Moayed
Menapace, Willi
Siarohin, Aliaksandr
Skorokhodov, Ivan
Canberk, Alper
Lee, Kwot Sin
Ordonez, Vicente
Tulyakov, Sergey
author_facet Haji-Ali, Moayed
Menapace, Willi
Siarohin, Aliaksandr
Skorokhodov, Ivan
Canberk, Alper
Lee, Kwot Sin
Ordonez, Vicente
Tulyakov, Sergey
contents We propose AV-Link, a unified framework for Video-to-Audio (A2V) and Audio-to-Video (A2V) generation that leverages the activations of frozen video and audio diffusion models for temporally-aligned cross-modal conditioning. The key to our framework is a Fusion Block that facilitates bidirectional information exchange between video and audio diffusion models through temporally-aligned self attention operations. Unlike prior work that uses dedicated models for A2V and V2A tasks and relies on pretrained feature extractors, AV-Link achieves both tasks in a single self-contained framework, directly leveraging features obtained by the complementary modality (i.e. video features to generate audio, or audio features to generate video). Extensive automatic and subjective evaluations demonstrate that our method achieves a substantial improvement in audio-video synchronization, outperforming more expensive baselines such as the MovieGen video-to-audio model.
format Preprint
id arxiv_https___arxiv_org_abs_2412_15191
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation
Haji-Ali, Moayed
Menapace, Willi
Siarohin, Aliaksandr
Skorokhodov, Ivan
Canberk, Alper
Lee, Kwot Sin
Ordonez, Vicente
Tulyakov, Sergey
Computer Vision and Pattern Recognition
Machine Learning
Sound
Audio and Speech Processing
We propose AV-Link, a unified framework for Video-to-Audio (A2V) and Audio-to-Video (A2V) generation that leverages the activations of frozen video and audio diffusion models for temporally-aligned cross-modal conditioning. The key to our framework is a Fusion Block that facilitates bidirectional information exchange between video and audio diffusion models through temporally-aligned self attention operations. Unlike prior work that uses dedicated models for A2V and V2A tasks and relies on pretrained feature extractors, AV-Link achieves both tasks in a single self-contained framework, directly leveraging features obtained by the complementary modality (i.e. video features to generate audio, or audio features to generate video). Extensive automatic and subjective evaluations demonstrate that our method achieves a substantial improvement in audio-video synchronization, outperforming more expensive baselines such as the MovieGen video-to-audio model.
title AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation
topic Computer Vision and Pattern Recognition
Machine Learning
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2412.15191