Images that Sound: Composing Images and Sounds on a Single Canvas

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Ziyang, Geng, Daniel, Owens, Andrew
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929698517286912
author Chen, Ziyang
Geng, Daniel
Owens, Andrew
author_facet Chen, Ziyang
Geng, Daniel
Owens, Andrew
contents Spectrograms are 2D representations of sound that look very different from the images found in our visual world. And natural images, when played as spectrograms, make unnatural sounds. In this paper, we show that it is possible to synthesize spectrograms that simultaneously look like natural images and sound like natural audio. We call these visual spectrograms images that sound. Our approach is simple and zero-shot, and it leverages pre-trained text-to-image and text-to-spectrogram diffusion models that operate in a shared latent space. During the reverse process, we denoise noisy latents with both the audio and image diffusion models in parallel, resulting in a sample that is likely under both models. Through quantitative evaluations and perceptual studies, we find that our method successfully generates spectrograms that align with a desired audio prompt while also taking the visual appearance of a desired image prompt. Please see our project page for video results: https://ificl.github.io/images-that-sound/
format Preprint
id arxiv_https___arxiv_org_abs_2405_12221
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Images that Sound: Composing Images and Sounds on a Single Canvas
Chen, Ziyang
Geng, Daniel
Owens, Andrew
Computer Vision and Pattern Recognition
Machine Learning
Multimedia
Sound
Audio and Speech Processing
Spectrograms are 2D representations of sound that look very different from the images found in our visual world. And natural images, when played as spectrograms, make unnatural sounds. In this paper, we show that it is possible to synthesize spectrograms that simultaneously look like natural images and sound like natural audio. We call these visual spectrograms images that sound. Our approach is simple and zero-shot, and it leverages pre-trained text-to-image and text-to-spectrogram diffusion models that operate in a shared latent space. During the reverse process, we denoise noisy latents with both the audio and image diffusion models in parallel, resulting in a sample that is likely under both models. Through quantitative evaluations and perceptual studies, we find that our method successfully generates spectrograms that align with a desired audio prompt while also taking the visual appearance of a desired image prompt. Please see our project page for video results: https://ificl.github.io/images-that-sound/
title Images that Sound: Composing Images and Sounds on a Single Canvas
topic Computer Vision and Pattern Recognition
Machine Learning
Multimedia
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2405.12221