Video-Guided Text-to-Music Generation Using Public Domain Movie Collections

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Haven, Novack, Zachary, Xu, Weihan, McAuley, Julian, Dong, Hao-Wen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908426512105472
author Kim, Haven
Novack, Zachary
Xu, Weihan
McAuley, Julian
Dong, Hao-Wen
author_facet Kim, Haven
Novack, Zachary
Xu, Weihan
McAuley, Julian
Dong, Hao-Wen
contents Despite recent advancements in music generation systems, their application in film production remains limited, as they struggle to capture the nuances of real-world filmmaking, where filmmakers consider multiple factors-such as visual content, dialogue, and emotional tone-when selecting or composing music for a scene. This limitation primarily stems from the absence of comprehensive datasets that integrate these elements. To address this gap, we introduce Open Screen Soundtrack Library (OSSL), a dataset consisting of movie clips from public domain films, totaling approximately 36.5 hours, paired with high-quality soundtracks and human-annotated mood information. To demonstrate the effectiveness of our dataset in improving the performance of pre-trained models on film music generation tasks, we introduce a new video adapter that enhances an autoregressive transformer-based text-to-music model by adding video-based conditioning. Our experimental results demonstrate that our proposed approach effectively enhances MusicGen-Medium in terms of both objective measures of distributional and paired fidelity, and subjective compatibility in mood and genre. To facilitate reproducibility and foster future work, we publicly release the dataset, code, and demo.
format Preprint
id arxiv_https___arxiv_org_abs_2506_12573
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Video-Guided Text-to-Music Generation Using Public Domain Movie Collections
Kim, Haven
Novack, Zachary
Xu, Weihan
McAuley, Julian
Dong, Hao-Wen
Sound
Multimedia
Audio and Speech Processing
Despite recent advancements in music generation systems, their application in film production remains limited, as they struggle to capture the nuances of real-world filmmaking, where filmmakers consider multiple factors-such as visual content, dialogue, and emotional tone-when selecting or composing music for a scene. This limitation primarily stems from the absence of comprehensive datasets that integrate these elements. To address this gap, we introduce Open Screen Soundtrack Library (OSSL), a dataset consisting of movie clips from public domain films, totaling approximately 36.5 hours, paired with high-quality soundtracks and human-annotated mood information. To demonstrate the effectiveness of our dataset in improving the performance of pre-trained models on film music generation tasks, we introduce a new video adapter that enhances an autoregressive transformer-based text-to-music model by adding video-based conditioning. Our experimental results demonstrate that our proposed approach effectively enhances MusicGen-Medium in terms of both objective measures of distributional and paired fidelity, and subjective compatibility in mood and genre. To facilitate reproducibility and foster future work, we publicly release the dataset, code, and demo.
title Video-Guided Text-to-Music Generation Using Public Domain Movie Collections
topic Sound
Multimedia
Audio and Speech Processing
url https://arxiv.org/abs/2506.12573