Scaling Up Audio-Synchronized Visual Animation: An Efficient Training Paradigm
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866908479777669120 |
|---|---|
| author | Zhang, Lin Cai, Zefan Zhou, Yufan Mo, Shentong Lin, Jinhong Wu, Cheng-En Wei, Yibing Zhang, Yijing Zhang, Ruiyi Xiao, Wen Sun, Tong Hu, Junjie Morgado, Pedro |
| author_facet | Zhang, Lin Cai, Zefan Zhou, Yufan Mo, Shentong Lin, Jinhong Wu, Cheng-En Wei, Yibing Zhang, Yijing Zhang, Ruiyi Xiao, Wen Sun, Tong Hu, Junjie Morgado, Pedro |
| contents | Recent advances in audio-synchronized visual animation enable control of video content using audios from specific classes. However, existing methods rely heavily on expensive manual curation of high-quality, class-specific training videos, posing challenges to scaling up to diverse audio-video classes in the open world. In this work, we propose an efficient two-stage training paradigm to scale up audio-synchronized visual animation using abundant but noisy videos. In stage one, we automatically curate large-scale videos for pretraining, allowing the model to learn diverse but imperfect audio-video alignments. In stage two, we finetune the model on manually curated high-quality examples, but only at a small scale, significantly reducing the required human effort. We further enhance synchronization by allowing each frame to access rich audio context via multi-feature conditioning and window attention. To efficiently train the model, we leverage pretrained text-to-video generator and audio encoders, introducing only 1.9\% additional trainable parameters to learn audio-conditioning capability without compromising the generator's prior knowledge. For evaluation, we introduce AVSync48, a benchmark with videos from 48 classes, which is 3$\times$ more diverse than previous benchmarks. Extensive experiments show that our method significantly reduces reliance on manual curation by over 10$\times$, while generalizing to many open classes. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_03955 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Scaling Up Audio-Synchronized Visual Animation: An Efficient Training Paradigm Zhang, Lin Cai, Zefan Zhou, Yufan Mo, Shentong Lin, Jinhong Wu, Cheng-En Wei, Yibing Zhang, Yijing Zhang, Ruiyi Xiao, Wen Sun, Tong Hu, Junjie Morgado, Pedro Computer Vision and Pattern Recognition Recent advances in audio-synchronized visual animation enable control of video content using audios from specific classes. However, existing methods rely heavily on expensive manual curation of high-quality, class-specific training videos, posing challenges to scaling up to diverse audio-video classes in the open world. In this work, we propose an efficient two-stage training paradigm to scale up audio-synchronized visual animation using abundant but noisy videos. In stage one, we automatically curate large-scale videos for pretraining, allowing the model to learn diverse but imperfect audio-video alignments. In stage two, we finetune the model on manually curated high-quality examples, but only at a small scale, significantly reducing the required human effort. We further enhance synchronization by allowing each frame to access rich audio context via multi-feature conditioning and window attention. To efficiently train the model, we leverage pretrained text-to-video generator and audio encoders, introducing only 1.9\% additional trainable parameters to learn audio-conditioning capability without compromising the generator's prior knowledge. For evaluation, we introduce AVSync48, a benchmark with videos from 48 classes, which is 3$\times$ more diverse than previous benchmarks. Extensive experiments show that our method significantly reduces reliance on manual curation by over 10$\times$, while generalizing to many open classes. |
| title | Scaling Up Audio-Synchronized Visual Animation: An Efficient Training Paradigm |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2508.03955 |