MorphFader: Enabling Fine-grained Controllable Morphing with Text-to-Audio Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kamath, Purnima, Gupta, Chitralekha, Nanayakkara, Suranga
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911987426918400
author Kamath, Purnima
Gupta, Chitralekha
Nanayakkara, Suranga
author_facet Kamath, Purnima
Gupta, Chitralekha
Nanayakkara, Suranga
contents Sound morphing is the process of gradually and smoothly transforming one sound into another to generate novel and perceptually hybrid sounds that simultaneously resemble both. Recently, diffusion-based text-to-audio models have produced high-quality sounds using text prompts. However, granularly controlling the semantics of the sound, which is necessary for morphing, can be challenging using text. In this paper, we propose \textit{MorphFader}, a controllable method for morphing sounds generated by disparate prompts using text-to-audio models. By intercepting and interpolating the components of the cross-attention layers within the diffusion process, we can create smooth morphs between sounds generated by different text prompts. Using both objective metrics and perceptual listening tests, we demonstrate the ability of our method to granularly control the semantics in the sound and generate smooth morphs.
format Preprint
id arxiv_https___arxiv_org_abs_2408_07260
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MorphFader: Enabling Fine-grained Controllable Morphing with Text-to-Audio Models
Kamath, Purnima
Gupta, Chitralekha
Nanayakkara, Suranga
Audio and Speech Processing
Sound morphing is the process of gradually and smoothly transforming one sound into another to generate novel and perceptually hybrid sounds that simultaneously resemble both. Recently, diffusion-based text-to-audio models have produced high-quality sounds using text prompts. However, granularly controlling the semantics of the sound, which is necessary for morphing, can be challenging using text. In this paper, we propose \textit{MorphFader}, a controllable method for morphing sounds generated by disparate prompts using text-to-audio models. By intercepting and interpolating the components of the cross-attention layers within the diffusion process, we can create smooth morphs between sounds generated by different text prompts. Using both objective metrics and perceptual listening tests, we demonstrate the ability of our method to granularly control the semantics in the sound and generate smooth morphs.
title MorphFader: Enabling Fine-grained Controllable Morphing with Text-to-Audio Models
topic Audio and Speech Processing
url https://arxiv.org/abs/2408.07260