AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909532861497344 |
|---|---|
| author | Haji-Ali, Moayed Menapace, Willi Siarohin, Aliaksandr Skorokhodov, Ivan Canberk, Alper Lee, Kwot Sin Ordonez, Vicente Tulyakov, Sergey |
| author_facet | Haji-Ali, Moayed Menapace, Willi Siarohin, Aliaksandr Skorokhodov, Ivan Canberk, Alper Lee, Kwot Sin Ordonez, Vicente Tulyakov, Sergey |
| contents | We propose AV-Link, a unified framework for Video-to-Audio (A2V) and Audio-to-Video (A2V) generation that leverages the activations of frozen video and audio diffusion models for temporally-aligned cross-modal conditioning. The key to our framework is a Fusion Block that facilitates bidirectional information exchange between video and audio diffusion models through temporally-aligned self attention operations. Unlike prior work that uses dedicated models for A2V and V2A tasks and relies on pretrained feature extractors, AV-Link achieves both tasks in a single self-contained framework, directly leveraging features obtained by the complementary modality (i.e. video features to generate audio, or audio features to generate video). Extensive automatic and subjective evaluations demonstrate that our method achieves a substantial improvement in audio-video synchronization, outperforming more expensive baselines such as the MovieGen video-to-audio model. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_15191 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Haji-Ali, Moayed Menapace, Willi Siarohin, Aliaksandr Skorokhodov, Ivan Canberk, Alper Lee, Kwot Sin Ordonez, Vicente Tulyakov, Sergey Computer Vision and Pattern Recognition Machine Learning Sound Audio and Speech Processing We propose AV-Link, a unified framework for Video-to-Audio (A2V) and Audio-to-Video (A2V) generation that leverages the activations of frozen video and audio diffusion models for temporally-aligned cross-modal conditioning. The key to our framework is a Fusion Block that facilitates bidirectional information exchange between video and audio diffusion models through temporally-aligned self attention operations. Unlike prior work that uses dedicated models for A2V and V2A tasks and relies on pretrained feature extractors, AV-Link achieves both tasks in a single self-contained framework, directly leveraging features obtained by the complementary modality (i.e. video features to generate audio, or audio features to generate video). Extensive automatic and subjective evaluations demonstrate that our method achieves a substantial improvement in audio-video synchronization, outperforming more expensive baselines such as the MovieGen video-to-audio model. |
| title | AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation |
| topic | Computer Vision and Pattern Recognition Machine Learning Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2412.15191 |