Model-Guided Dual-Role Alignment for High-Fidelity Open-Domain Video-to-Audio Generation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917047848402944 |
|---|---|
| author | Zhang, Kang Pham, Trung X. Lee, Suyeon Niu, Axi Senocak, Arda Chung, Joon Son |
| author_facet | Zhang, Kang Pham, Trung X. Lee, Suyeon Niu, Axi Senocak, Arda Chung, Joon Son |
| contents | We present MGAudio, a novel flow-based framework for open-domain video-to-audio generation, which introduces model-guided dual-role alignment as a central design principle. Unlike prior approaches that rely on classifier-based or classifier-free guidance, MGAudio enables the generative model to guide itself through a dedicated training objective designed for video-conditioned audio generation. The framework integrates three main components: (1) a scalable flow-based Transformer model, (2) a dual-role alignment mechanism where the audio-visual encoder serves both as a conditioning module and as a feature aligner to improve generation quality, and (3) a model-guided objective that enhances cross-modal coherence and audio realism. MGAudio achieves state-of-the-art performance on VGGSound, reducing FAD to 0.40, substantially surpassing the best classifier-free guidance baselines, and consistently outperforms existing methods across FD, IS, and alignment metrics. It also generalizes well to the challenging UnAV-100 benchmark. These results highlight model-guided dual-role alignment as a powerful and scalable paradigm for conditional video-to-audio generation. Code is available at: https://github.com/pantheon5100/mgaudio |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_24103 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Model-Guided Dual-Role Alignment for High-Fidelity Open-Domain Video-to-Audio Generation Zhang, Kang Pham, Trung X. Lee, Suyeon Niu, Axi Senocak, Arda Chung, Joon Son Sound Artificial Intelligence Multimedia Audio and Speech Processing We present MGAudio, a novel flow-based framework for open-domain video-to-audio generation, which introduces model-guided dual-role alignment as a central design principle. Unlike prior approaches that rely on classifier-based or classifier-free guidance, MGAudio enables the generative model to guide itself through a dedicated training objective designed for video-conditioned audio generation. The framework integrates three main components: (1) a scalable flow-based Transformer model, (2) a dual-role alignment mechanism where the audio-visual encoder serves both as a conditioning module and as a feature aligner to improve generation quality, and (3) a model-guided objective that enhances cross-modal coherence and audio realism. MGAudio achieves state-of-the-art performance on VGGSound, reducing FAD to 0.40, substantially surpassing the best classifier-free guidance baselines, and consistently outperforms existing methods across FD, IS, and alignment metrics. It also generalizes well to the challenging UnAV-100 benchmark. These results highlight model-guided dual-role alignment as a powerful and scalable paradigm for conditional video-to-audio generation. Code is available at: https://github.com/pantheon5100/mgaudio |
| title | Model-Guided Dual-Role Alignment for High-Fidelity Open-Domain Video-to-Audio Generation |
| topic | Sound Artificial Intelligence Multimedia Audio and Speech Processing |
| url | https://arxiv.org/abs/2510.24103 |