SAVGBench: Benchmarking Spatially Aligned Audio-Video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shimada, Kazuki, Simon, Christian, Shibuya, Takashi, Takahashi, Shusuke, Mitsufuji, Yuki
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917246123638784
author Shimada, Kazuki
Simon, Christian
Shibuya, Takashi
Takahashi, Shusuke
Mitsufuji, Yuki
author_facet Shimada, Kazuki
Simon, Christian
Shibuya, Takashi
Takahashi, Shusuke
Mitsufuji, Yuki
contents This work addresses the lack of multimodal generative models capable of producing high-quality videos with spatially aligned audio. While recent advancements in generative models have been successful in video generation, they often overlook the spatial alignment between audio and visuals, which is essential for immersive experiences. To tackle this problem, we establish a new research direction in benchmarking the Spatially Aligned Audio-Video Generation (SAVG) task. We introduce a spatially aligned audio-visual dataset, whose audio and video data are curated based on whether sound events are onscreen or not. We also propose a new alignment metric that aims to evaluate the spatial alignment between audio and video. Then, using the dataset and metric, we benchmark two types of baseline methods: one is based on a joint audio-video generation model, and the other is a two-stage method that combines a video generation model and a video-to-audio generation model. Our experimental results demonstrate that gaps exist between the baseline methods and the ground truth in terms of video and audio quality, as well as spatial alignment between the two modalities.
format Preprint
id arxiv_https___arxiv_org_abs_2412_13462
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SAVGBench: Benchmarking Spatially Aligned Audio-Video Generation
Shimada, Kazuki
Simon, Christian
Shibuya, Takashi
Takahashi, Shusuke
Mitsufuji, Yuki
Sound
Multimedia
Audio and Speech Processing
This work addresses the lack of multimodal generative models capable of producing high-quality videos with spatially aligned audio. While recent advancements in generative models have been successful in video generation, they often overlook the spatial alignment between audio and visuals, which is essential for immersive experiences. To tackle this problem, we establish a new research direction in benchmarking the Spatially Aligned Audio-Video Generation (SAVG) task. We introduce a spatially aligned audio-visual dataset, whose audio and video data are curated based on whether sound events are onscreen or not. We also propose a new alignment metric that aims to evaluate the spatial alignment between audio and video. Then, using the dataset and metric, we benchmark two types of baseline methods: one is based on a joint audio-video generation model, and the other is a two-stage method that combines a video generation model and a video-to-audio generation model. Our experimental results demonstrate that gaps exist between the baseline methods and the ground truth in terms of video and audio quality, as well as spatial alignment between the two modalities.
title SAVGBench: Benchmarking Spatially Aligned Audio-Video Generation
topic Sound
Multimedia
Audio and Speech Processing
url https://arxiv.org/abs/2412.13462