Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2512.02492 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912742860914688 |
|---|---|
| author | Chen, Jiahui Wang, Weida Shi, Runhua Yang, Huan Ding, Chaofan Chen, Zihao |
| author_facet | Chen, Jiahui Wang, Weida Shi, Runhua Yang, Huan Ding, Chaofan Chen, Zihao |
| contents | While diffusion model for audio-driven avatar video generation have achieved notable process in synthesizing long sequences with natural audio-visual synchronization and identity consistency, the generation of music-performance videos with camera motions remains largely unexplored. We present YingVideo-MV, the first cascaded framework for music-driven long-video generation. Our approach integrates audio semantic analysis, an interpretable shot planning module (MV-Director), temporal-aware diffusion Transformer architectures, and long-sequence consistency modeling to enable automatic synthesis of high-quality music performance videos from audio signals. We construct a large-scale Music-in-the-Wild Dataset by collecting web data to support the achievement of diverse, high-quality results. Observing that existing long-video generation methods lack explicit camera motion control, we introduce a camera adapter module that embeds camera poses into latent noise. To enhance continulity between clips during long-sequence inference, we further propose a time-aware dynamic window range strategy that adaptively adjust denoising ranges based on audio embedding. Comprehensive benchmark tests demonstrate that YingVideo-MV achieves outstanding performance in generating coherent and expressive music videos, and enables precise music-motion-camera synchronization. More videos are available in our project page: https://giantailab.github.io/YingVideo-MV/ . |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_02492 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | YingVideo-MV: Music-Driven Multi-Stage Video Generation Chen, Jiahui Wang, Weida Shi, Runhua Yang, Huan Ding, Chaofan Chen, Zihao Computer Vision and Pattern Recognition While diffusion model for audio-driven avatar video generation have achieved notable process in synthesizing long sequences with natural audio-visual synchronization and identity consistency, the generation of music-performance videos with camera motions remains largely unexplored. We present YingVideo-MV, the first cascaded framework for music-driven long-video generation. Our approach integrates audio semantic analysis, an interpretable shot planning module (MV-Director), temporal-aware diffusion Transformer architectures, and long-sequence consistency modeling to enable automatic synthesis of high-quality music performance videos from audio signals. We construct a large-scale Music-in-the-Wild Dataset by collecting web data to support the achievement of diverse, high-quality results. Observing that existing long-video generation methods lack explicit camera motion control, we introduce a camera adapter module that embeds camera poses into latent noise. To enhance continulity between clips during long-sequence inference, we further propose a time-aware dynamic window range strategy that adaptively adjust denoising ranges based on audio embedding. Comprehensive benchmark tests demonstrate that YingVideo-MV achieves outstanding performance in generating coherent and expressive music videos, and enables precise music-motion-camera synchronization. More videos are available in our project page: https://giantailab.github.io/YingVideo-MV/ . |
| title | YingVideo-MV: Music-Driven Multi-Stage Video Generation |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2512.02492 |