Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Hayakawa, Akio, Ishii, Masato, Shibuya, Takashi, Mitsufuji, Yuki
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916992275972096
author Hayakawa, Akio
Ishii, Masato
Shibuya, Takashi
Mitsufuji, Yuki
author_facet Hayakawa, Akio
Ishii, Masato
Shibuya, Takashi
Mitsufuji, Yuki
contents We propose a step-by-step video-to-audio (V2A) generation method for finer controllability over the generation process and more realistic audio synthesis. Inspired by traditional Foley workflows, our approach aims to comprehensively capture all sound events induced by a video through the incremental generation of missing sound events. To avoid the need for costly multi-reference video-audio datasets, each generation step is formulated as a negatively guided V2A process that discourages duplication of existing sounds. The guidance model is trained by finetuning a pre-trained V2A model on audio pairs from adjacent segments of the same video, allowing training with standard single-reference audiovisual datasets that are easily accessible. Objective and subjective evaluations demonstrate that our method enhances the separability of generated sounds at each step and improves the overall quality of the final composite audio, outperforming existing baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2506_20995
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance
Hayakawa, Akio
Ishii, Masato
Shibuya, Takashi
Mitsufuji, Yuki
Computer Vision and Pattern Recognition
Machine Learning
Sound
Audio and Speech Processing
We propose a step-by-step video-to-audio (V2A) generation method for finer controllability over the generation process and more realistic audio synthesis. Inspired by traditional Foley workflows, our approach aims to comprehensively capture all sound events induced by a video through the incremental generation of missing sound events. To avoid the need for costly multi-reference video-audio datasets, each generation step is formulated as a negatively guided V2A process that discourages duplication of existing sounds. The guidance model is trained by finetuning a pre-trained V2A model on audio pairs from adjacent segments of the same video, allowing training with standard single-reference audiovisual datasets that are easily accessible. Objective and subjective evaluations demonstrate that our method enhances the separability of generated sounds at each step and improves the overall quality of the final composite audio, outperforming existing baselines.
title Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance
topic Computer Vision and Pattern Recognition
Machine Learning
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2506.20995