StereoFoley: Object-Aware Stereo Audio Generation from Video

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Karchkhadze, Tornike, Chen, Kuan-Lin, Heydari, Mojtaba, Henzel, Robert, Toso, Alessandro, Souden, Mehrez, Atkins, Joshua
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913041311858688
author Karchkhadze, Tornike
Chen, Kuan-Lin
Heydari, Mojtaba
Henzel, Robert
Toso, Alessandro
Souden, Mehrez
Atkins, Joshua
author_facet Karchkhadze, Tornike
Chen, Kuan-Lin
Heydari, Mojtaba
Henzel, Robert
Toso, Alessandro
Souden, Mehrez
Atkins, Joshua
contents We present StereoFoley, a video-to-audio generation framework that produces semantically aligned, temporally synchronized, and spatially accurate stereo sound at 48 kHz. While recent generative video-to-audio models achieve strong semantic and temporal fidelity, they largely remain limited to mono or fail to deliver object-aware stereo imaging, constrained by the lack of professionally mixed, spatially accurate video-to-audio datasets. First, we develop a base model that generates stereo audio from video, achieving performance on par with state-of-the-art V2A models in both semantic accuracy and synchronization. Next, to overcome dataset limitations, we introduce a synthetic data generation pipeline that combines video analysis, object tracking, and audio synthesis with dynamic panning and distance-based loudness controls, enabling spatially accurate object-aware sound. Finally, we fine-tune the base model on this synthetic dataset, yielding clear object-audio correspondence. Since no established metrics exist, we introduce a stereo object-awareness metric and report it alongside a human listening study; the two evaluations exhibit consistent trends. This work establishes the first end-to-end framework for stereo object-aware video-to-audio generation, addressing a critical gap in the field.
format Preprint
id arxiv_https___arxiv_org_abs_2509_18272
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle StereoFoley: Object-Aware Stereo Audio Generation from Video
Karchkhadze, Tornike
Chen, Kuan-Lin
Heydari, Mojtaba
Henzel, Robert
Toso, Alessandro
Souden, Mehrez
Atkins, Joshua
Sound
Multimedia
Audio and Speech Processing
We present StereoFoley, a video-to-audio generation framework that produces semantically aligned, temporally synchronized, and spatially accurate stereo sound at 48 kHz. While recent generative video-to-audio models achieve strong semantic and temporal fidelity, they largely remain limited to mono or fail to deliver object-aware stereo imaging, constrained by the lack of professionally mixed, spatially accurate video-to-audio datasets. First, we develop a base model that generates stereo audio from video, achieving performance on par with state-of-the-art V2A models in both semantic accuracy and synchronization. Next, to overcome dataset limitations, we introduce a synthetic data generation pipeline that combines video analysis, object tracking, and audio synthesis with dynamic panning and distance-based loudness controls, enabling spatially accurate object-aware sound. Finally, we fine-tune the base model on this synthetic dataset, yielding clear object-audio correspondence. Since no established metrics exist, we introduce a stereo object-awareness metric and report it alongside a human listening study; the two evaluations exhibit consistent trends. This work establishes the first end-to-end framework for stereo object-aware video-to-audio generation, addressing a critical gap in the field.
title StereoFoley: Object-Aware Stereo Audio Generation from Video
topic Sound
Multimedia
Audio and Speech Processing
url https://arxiv.org/abs/2509.18272