Diff-BGM: A Diffusion Model for Video Background Music Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Sizhe, Qin, Yiming, Zheng, Minghang, Jin, Xin, Liu, Yang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913357007683584
author Li, Sizhe
Qin, Yiming
Zheng, Minghang
Jin, Xin
Liu, Yang
author_facet Li, Sizhe
Qin, Yiming
Zheng, Minghang
Jin, Xin
Liu, Yang
contents When editing a video, a piece of attractive background music is indispensable. However, video background music generation tasks face several challenges, for example, the lack of suitable training datasets, and the difficulties in flexibly controlling the music generation process and sequentially aligning the video and music. In this work, we first propose a high-quality music-video dataset BGM909 with detailed annotation and shot detection to provide multi-modal information about the video and music. We then present evaluation metrics to assess music quality, including music diversity and alignment between music and video with retrieval precision metrics. Finally, we propose the Diff-BGM framework to automatically generate the background music for a given video, which uses different signals to control different aspects of the music during the generation process, i.e., uses dynamic video features to control music rhythm and semantic features to control the melody and atmosphere. We propose to align the video and music sequentially by introducing a segment-aware cross-attention layer. Experiments verify the effectiveness of our proposed method. The code and models are available at https://github.com/sizhelee/Diff-BGM.
format Preprint
id arxiv_https___arxiv_org_abs_2405_11913
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Diff-BGM: A Diffusion Model for Video Background Music Generation
Li, Sizhe
Qin, Yiming
Zheng, Minghang
Jin, Xin
Liu, Yang
Computer Vision and Pattern Recognition
When editing a video, a piece of attractive background music is indispensable. However, video background music generation tasks face several challenges, for example, the lack of suitable training datasets, and the difficulties in flexibly controlling the music generation process and sequentially aligning the video and music. In this work, we first propose a high-quality music-video dataset BGM909 with detailed annotation and shot detection to provide multi-modal information about the video and music. We then present evaluation metrics to assess music quality, including music diversity and alignment between music and video with retrieval precision metrics. Finally, we propose the Diff-BGM framework to automatically generate the background music for a given video, which uses different signals to control different aspects of the music during the generation process, i.e., uses dynamic video features to control music rhythm and semantic features to control the melody and atmosphere. We propose to align the video and music sequentially by introducing a segment-aware cross-attention layer. Experiments verify the effectiveness of our proposed method. The code and models are available at https://github.com/sizhelee/Diff-BGM.
title Diff-BGM: A Diffusion Model for Video Background Music Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.11913