Scaling Up Audio-Synchronized Visual Animation: An Efficient Training Paradigm

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Lin, Cai, Zefan, Zhou, Yufan, Mo, Shentong, Lin, Jinhong, Wu, Cheng-En, Wei, Yibing, Zhang, Yijing, Zhang, Ruiyi, Xiao, Wen, Sun, Tong, Hu, Junjie, Morgado, Pedro
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908479777669120
author Zhang, Lin
Cai, Zefan
Zhou, Yufan
Mo, Shentong
Lin, Jinhong
Wu, Cheng-En
Wei, Yibing
Zhang, Yijing
Zhang, Ruiyi
Xiao, Wen
Sun, Tong
Hu, Junjie
Morgado, Pedro
author_facet Zhang, Lin
Cai, Zefan
Zhou, Yufan
Mo, Shentong
Lin, Jinhong
Wu, Cheng-En
Wei, Yibing
Zhang, Yijing
Zhang, Ruiyi
Xiao, Wen
Sun, Tong
Hu, Junjie
Morgado, Pedro
contents Recent advances in audio-synchronized visual animation enable control of video content using audios from specific classes. However, existing methods rely heavily on expensive manual curation of high-quality, class-specific training videos, posing challenges to scaling up to diverse audio-video classes in the open world. In this work, we propose an efficient two-stage training paradigm to scale up audio-synchronized visual animation using abundant but noisy videos. In stage one, we automatically curate large-scale videos for pretraining, allowing the model to learn diverse but imperfect audio-video alignments. In stage two, we finetune the model on manually curated high-quality examples, but only at a small scale, significantly reducing the required human effort. We further enhance synchronization by allowing each frame to access rich audio context via multi-feature conditioning and window attention. To efficiently train the model, we leverage pretrained text-to-video generator and audio encoders, introducing only 1.9\% additional trainable parameters to learn audio-conditioning capability without compromising the generator's prior knowledge. For evaluation, we introduce AVSync48, a benchmark with videos from 48 classes, which is 3$\times$ more diverse than previous benchmarks. Extensive experiments show that our method significantly reduces reliance on manual curation by over 10$\times$, while generalizing to many open classes.
format Preprint
id arxiv_https___arxiv_org_abs_2508_03955
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Scaling Up Audio-Synchronized Visual Animation: An Efficient Training Paradigm
Zhang, Lin
Cai, Zefan
Zhou, Yufan
Mo, Shentong
Lin, Jinhong
Wu, Cheng-En
Wei, Yibing
Zhang, Yijing
Zhang, Ruiyi
Xiao, Wen
Sun, Tong
Hu, Junjie
Morgado, Pedro
Computer Vision and Pattern Recognition
Recent advances in audio-synchronized visual animation enable control of video content using audios from specific classes. However, existing methods rely heavily on expensive manual curation of high-quality, class-specific training videos, posing challenges to scaling up to diverse audio-video classes in the open world. In this work, we propose an efficient two-stage training paradigm to scale up audio-synchronized visual animation using abundant but noisy videos. In stage one, we automatically curate large-scale videos for pretraining, allowing the model to learn diverse but imperfect audio-video alignments. In stage two, we finetune the model on manually curated high-quality examples, but only at a small scale, significantly reducing the required human effort. We further enhance synchronization by allowing each frame to access rich audio context via multi-feature conditioning and window attention. To efficiently train the model, we leverage pretrained text-to-video generator and audio encoders, introducing only 1.9\% additional trainable parameters to learn audio-conditioning capability without compromising the generator's prior knowledge. For evaluation, we introduce AVSync48, a benchmark with videos from 48 classes, which is 3$\times$ more diverse than previous benchmarks. Extensive experiments show that our method significantly reduces reliance on manual curation by over 10$\times$, while generalizing to many open classes.
title Scaling Up Audio-Synchronized Visual Animation: An Efficient Training Paradigm
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.03955