Video Echoed in Music: Semantic, Temporal, and Rhythmic Alignment for Video-to-Music Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tong, Xinyi, Zhu, Yiran, Chen, Jishang, Zhan, Chunru, Wang, Tianle, Zhang, Sirui, Liu, Nian, Ge, Tiezheng, Xu, Duo, Jin, Xin, Yu, Feng, Zhu, Song-Chun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918245308563456
author Tong, Xinyi
Zhu, Yiran
Chen, Jishang
Zhan, Chunru
Wang, Tianle
Zhang, Sirui
Liu, Nian
Ge, Tiezheng
Xu, Duo
Jin, Xin
Yu, Feng
Zhu, Song-Chun
author_facet Tong, Xinyi
Zhu, Yiran
Chen, Jishang
Zhan, Chunru
Wang, Tianle
Zhang, Sirui
Liu, Nian
Ge, Tiezheng
Xu, Duo
Jin, Xin
Yu, Feng
Zhu, Song-Chun
contents Video-to-Music generation seeks to generate musically appropriate background music that enhances audiovisual immersion for videos. However, current approaches suffer from two critical limitations: 1) incomplete representation of video details, leading to weak alignment, and 2) inadequate temporal and rhythmic correspondence, particularly in achieving precise beat synchronization. To address the challenges, we propose Video Echoed in Music (VeM), a latent music diffusion that generates high-quality soundtracks with semantic, temporal, and rhythmic alignment for input videos. To capture video details comprehensively, VeM employs a hierarchical video parsing that acts as a music conductor, orchestrating multi-level information across modalities. Modality-specific encoders, coupled with a storyboard-guided cross-attention mechanism (SG-CAtt), integrate semantic cues while maintaining temporal coherence through position and duration encoding. For rhythmic precision, the frame-level transition-beat aligner and adapter (TB-As) dynamically synchronize visual scene transitions with music beats. We further contribute a novel video-music paired dataset sourced from e-commerce advertisements and video-sharing platforms, which imposes stricter transition-beat synchronization requirements. Meanwhile, we introduce novel metrics tailored to the task. Experimental results demonstrate superiority, particularly in semantic relevance and rhythmic precision.
format Preprint
id arxiv_https___arxiv_org_abs_2511_09585
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Video Echoed in Music: Semantic, Temporal, and Rhythmic Alignment for Video-to-Music Generation
Tong, Xinyi
Zhu, Yiran
Chen, Jishang
Zhan, Chunru
Wang, Tianle
Zhang, Sirui
Liu, Nian
Ge, Tiezheng
Xu, Duo
Jin, Xin
Yu, Feng
Zhu, Song-Chun
Sound
Multimedia
Video-to-Music generation seeks to generate musically appropriate background music that enhances audiovisual immersion for videos. However, current approaches suffer from two critical limitations: 1) incomplete representation of video details, leading to weak alignment, and 2) inadequate temporal and rhythmic correspondence, particularly in achieving precise beat synchronization. To address the challenges, we propose Video Echoed in Music (VeM), a latent music diffusion that generates high-quality soundtracks with semantic, temporal, and rhythmic alignment for input videos. To capture video details comprehensively, VeM employs a hierarchical video parsing that acts as a music conductor, orchestrating multi-level information across modalities. Modality-specific encoders, coupled with a storyboard-guided cross-attention mechanism (SG-CAtt), integrate semantic cues while maintaining temporal coherence through position and duration encoding. For rhythmic precision, the frame-level transition-beat aligner and adapter (TB-As) dynamically synchronize visual scene transitions with music beats. We further contribute a novel video-music paired dataset sourced from e-commerce advertisements and video-sharing platforms, which imposes stricter transition-beat synchronization requirements. Meanwhile, we introduce novel metrics tailored to the task. Experimental results demonstrate superiority, particularly in semantic relevance and rhythmic precision.
title Video Echoed in Music: Semantic, Temporal, and Rhythmic Alignment for Video-to-Music Generation
topic Sound
Multimedia
url https://arxiv.org/abs/2511.09585