Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Petermann, Darius, Kalayeh, Mahdi M.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910778429276160
author Petermann, Darius
Kalayeh, Mahdi M.
author_facet Petermann, Darius
Kalayeh, Mahdi M.
contents Training audio-to-image generative models requires an abundance of diverse audio-visual pairs that are semantically aligned. Such data is almost always curated from in-the-wild videos, given the cross-modal semantic correspondence that is inherent to them. In this work, we hypothesize that insisting on the absolute need for ground truth audio-visual correspondence, is not only unnecessary, but also leads to severe restrictions in scale, quality, and diversity of the data, ultimately impairing its use in the modern generative models. That is, we propose a scalable image sonification framework where instances from a variety of high-quality yet disjoint uni-modal origins can be artificially paired through a retrieval process that is empowered by reasoning capabilities of modern vision-language models. To demonstrate the efficacy of this approach, we use our sonified images to train an audio-to-image generative model that performs competitively against state-of-the-art. Finally, through a series of ablation studies, we exhibit several intriguing auditory capabilities like semantic mixing and interpolation, loudness calibration and acoustic space modeling through reverberation that our model has implicitly developed to guide the image generation process.
format Preprint
id arxiv_https___arxiv_org_abs_2501_05413
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation
Petermann, Darius
Kalayeh, Mahdi M.
Sound
Computer Vision and Pattern Recognition
Graphics
Audio and Speech Processing
Training audio-to-image generative models requires an abundance of diverse audio-visual pairs that are semantically aligned. Such data is almost always curated from in-the-wild videos, given the cross-modal semantic correspondence that is inherent to them. In this work, we hypothesize that insisting on the absolute need for ground truth audio-visual correspondence, is not only unnecessary, but also leads to severe restrictions in scale, quality, and diversity of the data, ultimately impairing its use in the modern generative models. That is, we propose a scalable image sonification framework where instances from a variety of high-quality yet disjoint uni-modal origins can be artificially paired through a retrieval process that is empowered by reasoning capabilities of modern vision-language models. To demonstrate the efficacy of this approach, we use our sonified images to train an audio-to-image generative model that performs competitively against state-of-the-art. Finally, through a series of ablation studies, we exhibit several intriguing auditory capabilities like semantic mixing and interpolation, loudness calibration and acoustic space modeling through reverberation that our model has implicitly developed to guide the image generation process.
title Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation
topic Sound
Computer Vision and Pattern Recognition
Graphics
Audio and Speech Processing
url https://arxiv.org/abs/2501.05413