Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yaoyun, Xu, Xuenan, Wu, Mengyue
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913624859082752
author Zhang, Yaoyun
Xu, Xuenan
Wu, Mengyue
author_facet Zhang, Yaoyun
Xu, Xuenan
Wu, Mengyue
contents The video-to-audio (V2A) generation task has drawn attention in the field of multimedia due to the practicality in producing Foley sound. Semantic and temporal conditions are fed to the generation model to indicate sound events and temporal occurrence. Recent studies on synthesizing immersive and synchronized audio are faced with challenges on videos with moving visual presence. The temporal condition is not accurate enough while low-resolution semantic condition exacerbates the problem. To tackle these challenges, we propose Smooth-Foley, a V2A generative model taking semantic guidance from the textual label across the generation to enhance both semantic and temporal alignment in audio. Two adapters are trained to leverage pre-trained text-to-audio generation models. A frame adapter integrates high-resolution frame-wise video features while a temporal adapter integrates temporal conditions obtained from similarities of visual frames and textual labels. The incorporation of semantic guidance from textual labels achieves precise audio-video alignment. We conduct extensive quantitative and qualitative experiments. Results show that Smooth-Foley performs better than existing models on both continuous sound scenarios and general scenarios. With semantic guidance, the audio generated by Smooth-Foley exhibits higher quality and better adherence to physical laws.
format Preprint
id arxiv_https___arxiv_org_abs_2412_18157
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance
Zhang, Yaoyun
Xu, Xuenan
Wu, Mengyue
Sound
Artificial Intelligence
Audio and Speech Processing
The video-to-audio (V2A) generation task has drawn attention in the field of multimedia due to the practicality in producing Foley sound. Semantic and temporal conditions are fed to the generation model to indicate sound events and temporal occurrence. Recent studies on synthesizing immersive and synchronized audio are faced with challenges on videos with moving visual presence. The temporal condition is not accurate enough while low-resolution semantic condition exacerbates the problem. To tackle these challenges, we propose Smooth-Foley, a V2A generative model taking semantic guidance from the textual label across the generation to enhance both semantic and temporal alignment in audio. Two adapters are trained to leverage pre-trained text-to-audio generation models. A frame adapter integrates high-resolution frame-wise video features while a temporal adapter integrates temporal conditions obtained from similarities of visual frames and textual labels. The incorporation of semantic guidance from textual labels achieves precise audio-video alignment. We conduct extensive quantitative and qualitative experiments. Results show that Smooth-Foley performs better than existing models on both continuous sound scenarios and general scenarios. With semantic guidance, the audio generated by Smooth-Foley exhibits higher quality and better adherence to physical laws.
title Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2412.18157