Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Chang, Zhang, Haomin, Xia, Shiyu, Chen, Zihao, Ding, Chaofan, Yue, Xin, Chen, Huizhe, Di, Xinhan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909623456366592
author Liu, Chang
Zhang, Haomin
Xia, Shiyu
Chen, Zihao
Ding, Chaofan
Yue, Xin
Chen, Huizhe
Di, Xinhan
author_facet Liu, Chang
Zhang, Haomin
Xia, Shiyu
Chen, Zihao
Ding, Chaofan
Yue, Xin
Chen, Huizhe
Di, Xinhan
contents Generating high-quality piano audio from video requires precise synchronization between visual cues and musical output, ensuring accurate semantic and temporal alignment.However, existing evaluation datasets do not fully capture the intricate synchronization required for piano music generation. A comprehensive benchmark is essential for two primary reasons: (1) existing metrics fail to reflect the complexity of video-to-piano music interactions, and (2) a dedicated benchmark dataset can provide valuable insights to accelerate progress in high-quality piano music generation. To address these challenges, we introduce the CoP Benchmark Dataset-a fully open-sourced, multimodal benchmark designed specifically for video-guided piano music generation. The proposed Chain-of-Perform (CoP) benchmark offers several compelling features: (1) detailed multimodal annotations, enabling precise semantic and temporal alignment between video content and piano audio via step-by-step Chain-of-Perform guidance; (2) a versatile evaluation framework for rigorous assessment of both general-purpose and specialized video-to-piano generation tasks; and (3) full open-sourcing of the dataset, annotations, and evaluation protocols. The dataset is publicly available at https://github.com/acappemin/Video-to-Audio-and-Piano, with a continuously updated leaderboard to promote ongoing research in this domain.
format Preprint
id arxiv_https___arxiv_org_abs_2505_20038
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks
Liu, Chang
Zhang, Haomin
Xia, Shiyu
Chen, Zihao
Ding, Chaofan
Yue, Xin
Chen, Huizhe
Di, Xinhan
Sound
Computer Vision and Pattern Recognition
Audio and Speech Processing
Generating high-quality piano audio from video requires precise synchronization between visual cues and musical output, ensuring accurate semantic and temporal alignment.However, existing evaluation datasets do not fully capture the intricate synchronization required for piano music generation. A comprehensive benchmark is essential for two primary reasons: (1) existing metrics fail to reflect the complexity of video-to-piano music interactions, and (2) a dedicated benchmark dataset can provide valuable insights to accelerate progress in high-quality piano music generation. To address these challenges, we introduce the CoP Benchmark Dataset-a fully open-sourced, multimodal benchmark designed specifically for video-guided piano music generation. The proposed Chain-of-Perform (CoP) benchmark offers several compelling features: (1) detailed multimodal annotations, enabling precise semantic and temporal alignment between video content and piano audio via step-by-step Chain-of-Perform guidance; (2) a versatile evaluation framework for rigorous assessment of both general-purpose and specialized video-to-piano generation tasks; and (3) full open-sourcing of the dataset, annotations, and evaluation protocols. The dataset is publicly available at https://github.com/acappemin/Video-to-Audio-and-Piano, with a continuously updated leaderboard to promote ongoing research in this domain.
title Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks
topic Sound
Computer Vision and Pattern Recognition
Audio and Speech Processing
url https://arxiv.org/abs/2505.20038