SonicDiffusion: Audio-Driven Image Generation and Editing with Pretrained Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Biner, Burak Can, Sofian, Farrin Marouf, Karakaş, Umur Berkay, Ceylan, Duygu, Erdem, Erkut, Erdem, Aykut
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929333375860736
author Biner, Burak Can
Sofian, Farrin Marouf
Karakaş, Umur Berkay
Ceylan, Duygu
Erdem, Erkut
Erdem, Aykut
author_facet Biner, Burak Can
Sofian, Farrin Marouf
Karakaş, Umur Berkay
Ceylan, Duygu
Erdem, Erkut
Erdem, Aykut
contents We are witnessing a revolution in conditional image synthesis with the recent success of large scale text-to-image generation methods. This success also opens up new opportunities in controlling the generation and editing process using multi-modal input. While spatial control using cues such as depth, sketch, and other images has attracted a lot of research, we argue that another equally effective modality is audio since sound and sight are two main components of human perception. Hence, we propose a method to enable audio-conditioning in large scale image diffusion models. Our method first maps features obtained from audio clips to tokens that can be injected into the diffusion model in a fashion similar to text tokens. We introduce additional audio-image cross attention layers which we finetune while freezing the weights of the original layers of the diffusion model. In addition to audio conditioned image generation, our method can also be utilized in conjuction with diffusion based editing methods to enable audio conditioned image editing. We demonstrate our method on a wide range of audio and image datasets. We perform extensive comparisons with recent methods and show favorable performance.
format Preprint
id arxiv_https___arxiv_org_abs_2405_00878
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SonicDiffusion: Audio-Driven Image Generation and Editing with Pretrained Diffusion Models
Biner, Burak Can
Sofian, Farrin Marouf
Karakaş, Umur Berkay
Ceylan, Duygu
Erdem, Erkut
Erdem, Aykut
Computer Vision and Pattern Recognition
We are witnessing a revolution in conditional image synthesis with the recent success of large scale text-to-image generation methods. This success also opens up new opportunities in controlling the generation and editing process using multi-modal input. While spatial control using cues such as depth, sketch, and other images has attracted a lot of research, we argue that another equally effective modality is audio since sound and sight are two main components of human perception. Hence, we propose a method to enable audio-conditioning in large scale image diffusion models. Our method first maps features obtained from audio clips to tokens that can be injected into the diffusion model in a fashion similar to text tokens. We introduce additional audio-image cross attention layers which we finetune while freezing the weights of the original layers of the diffusion model. In addition to audio conditioned image generation, our method can also be utilized in conjuction with diffusion based editing methods to enable audio conditioned image editing. We demonstrate our method on a wide range of audio and image datasets. We perform extensive comparisons with recent methods and show favorable performance.
title SonicDiffusion: Audio-Driven Image Generation and Editing with Pretrained Diffusion Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.00878