Vision-to-Music Generation: A Survey

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Zhaokai, Bao, Chenxi, Zhuo, Le, Han, Jingrui, Yue, Yang, Tang, Yihong, Huang, Victor Shea-Jay, Liao, Yue
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913975328833536
author Wang, Zhaokai
Bao, Chenxi
Zhuo, Le
Han, Jingrui
Yue, Yang
Tang, Yihong
Huang, Victor Shea-Jay
Liao, Yue
author_facet Wang, Zhaokai
Bao, Chenxi
Zhuo, Le
Han, Jingrui
Yue, Yang
Tang, Yihong
Huang, Victor Shea-Jay
Liao, Yue
contents Vision-to-music Generation, including video-to-music and image-to-music tasks, is a significant branch of multimodal artificial intelligence demonstrating vast application prospects in fields such as film scoring, short video creation, and dance music synthesis. However, compared to the rapid development of modalities like text and images, research in vision-to-music is still in its preliminary stage due to its complex internal structure and the difficulty of modeling dynamic relationships with video. Existing surveys focus on general music generation without comprehensive discussion on vision-to-music. In this paper, we systematically review the research progress in the field of vision-to-music generation. We first analyze the technical characteristics and core challenges for three input types: general videos, human movement videos, and images, as well as two output types of symbolic music and audio music. We then summarize the existing methodologies on vision-to-music generation from the architecture perspective. A detailed review of common datasets and evaluation metrics is provided. Finally, we discuss current challenges and promising directions for future research. We hope our survey can inspire further innovation in vision-to-music generation and the broader field of multimodal generation in academic research and industrial applications. To follow latest works and foster further innovation in this field, we are continuously maintaining a GitHub repository at https://github.com/wzk1015/Awesome-Vision-to-Music-Generation.
format Preprint
id arxiv_https___arxiv_org_abs_2503_21254
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Vision-to-Music Generation: A Survey
Wang, Zhaokai
Bao, Chenxi
Zhuo, Le
Han, Jingrui
Yue, Yang
Tang, Yihong
Huang, Victor Shea-Jay
Liao, Yue
Computer Vision and Pattern Recognition
Artificial Intelligence
Multimedia
Sound
Audio and Speech Processing
Vision-to-music Generation, including video-to-music and image-to-music tasks, is a significant branch of multimodal artificial intelligence demonstrating vast application prospects in fields such as film scoring, short video creation, and dance music synthesis. However, compared to the rapid development of modalities like text and images, research in vision-to-music is still in its preliminary stage due to its complex internal structure and the difficulty of modeling dynamic relationships with video. Existing surveys focus on general music generation without comprehensive discussion on vision-to-music. In this paper, we systematically review the research progress in the field of vision-to-music generation. We first analyze the technical characteristics and core challenges for three input types: general videos, human movement videos, and images, as well as two output types of symbolic music and audio music. We then summarize the existing methodologies on vision-to-music generation from the architecture perspective. A detailed review of common datasets and evaluation metrics is provided. Finally, we discuss current challenges and promising directions for future research. We hope our survey can inspire further innovation in vision-to-music generation and the broader field of multimodal generation in academic research and industrial applications. To follow latest works and foster further innovation in this field, we are continuously maintaining a GitHub repository at https://github.com/wzk1015/Awesome-Vision-to-Music-Generation.
title Vision-to-Music Generation: A Survey
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Multimedia
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2503.21254