AutoMV: An Automatic Multi-Agent System for Music Video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tang, Xiaoxuan, Lei, Xinping, Zhu, Chaoran, Chen, Shiyun, Yuan, Ruibin, Li, Yizhi, Oh, Changjae, Zhang, Ge, Huang, Wenhao, Benetos, Emmanouil, Liu, Yang, Liu, Jiaheng, Ma, Yinghao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909959686455296
author Tang, Xiaoxuan
Lei, Xinping
Zhu, Chaoran
Chen, Shiyun
Yuan, Ruibin
Li, Yizhi
Oh, Changjae
Zhang, Ge
Huang, Wenhao
Benetos, Emmanouil
Liu, Yang
Liu, Jiaheng
Ma, Yinghao
author_facet Tang, Xiaoxuan
Lei, Xinping
Zhu, Chaoran
Chen, Shiyun
Yuan, Ruibin
Li, Yizhi
Oh, Changjae
Zhang, Ge
Huang, Wenhao
Benetos, Emmanouil
Liu, Yang
Liu, Jiaheng
Ma, Yinghao
contents Music-to-Video (M2V) generation for full-length songs faces significant challenges. Existing methods produce short, disjointed clips, failing to align visuals with musical structure, beats, or lyrics, and lack temporal consistency. We propose AutoMV, a multi-agent system that generates full music videos (MVs) directly from a song. AutoMV first applies music processing tools to extract musical attributes, such as structure, vocal tracks, and time-aligned lyrics, and constructs these features as contextual inputs for following agents. The screenwriter Agent and director Agent then use this information to design short script, define character profiles in a shared external bank, and specify camera instructions. Subsequently, these agents call the image generator for keyframes and different video generators for "story" or "singer" scenes. A Verifier Agent evaluates their output, enabling multi-agent collaboration to produce a coherent longform MV. To evaluate M2V generation, we further propose a benchmark with four high-level categories (Music Content, Technical, Post-production, Art) and twelve ine-grained criteria. This benchmark was applied to compare commercial products, AutoMV, and human-directed MVs with expert human raters: AutoMV outperforms current baselines significantly across all four categories, narrowing the gap to professional MVs. Finally, we investigate using large multimodal models as automatic MV judges; while promising, they still lag behind human expert, highlighting room for future work.
format Preprint
id arxiv_https___arxiv_org_abs_2512_12196
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AutoMV: An Automatic Multi-Agent System for Music Video Generation
Tang, Xiaoxuan
Lei, Xinping
Zhu, Chaoran
Chen, Shiyun
Yuan, Ruibin
Li, Yizhi
Oh, Changjae
Zhang, Ge
Huang, Wenhao
Benetos, Emmanouil
Liu, Yang
Liu, Jiaheng
Ma, Yinghao
Multimedia
Computer Vision and Pattern Recognition
Sound
Audio and Speech Processing
Music-to-Video (M2V) generation for full-length songs faces significant challenges. Existing methods produce short, disjointed clips, failing to align visuals with musical structure, beats, or lyrics, and lack temporal consistency. We propose AutoMV, a multi-agent system that generates full music videos (MVs) directly from a song. AutoMV first applies music processing tools to extract musical attributes, such as structure, vocal tracks, and time-aligned lyrics, and constructs these features as contextual inputs for following agents. The screenwriter Agent and director Agent then use this information to design short script, define character profiles in a shared external bank, and specify camera instructions. Subsequently, these agents call the image generator for keyframes and different video generators for "story" or "singer" scenes. A Verifier Agent evaluates their output, enabling multi-agent collaboration to produce a coherent longform MV. To evaluate M2V generation, we further propose a benchmark with four high-level categories (Music Content, Technical, Post-production, Art) and twelve ine-grained criteria. This benchmark was applied to compare commercial products, AutoMV, and human-directed MVs with expert human raters: AutoMV outperforms current baselines significantly across all four categories, narrowing the gap to professional MVs. Finally, we investigate using large multimodal models as automatic MV judges; while promising, they still lag behind human expert, highlighting room for future work.
title AutoMV: An Automatic Multi-Agent System for Music Video Generation
topic Multimedia
Computer Vision and Pattern Recognition
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2512.12196