Prompt-guided Precise Audio Editing with Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Manjie, Li, Chenxing, zhang, Duzhen, Su, Dan, Liang, Wei, Yu, Dong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910475392909312
author Xu, Manjie
Li, Chenxing
zhang, Duzhen
Su, Dan
Liang, Wei
Yu, Dong
author_facet Xu, Manjie
Li, Chenxing
zhang, Duzhen
Su, Dan
Liang, Wei
Yu, Dong
contents Audio editing involves the arbitrary manipulation of audio content through precise control. Although text-guided diffusion models have made significant advancements in text-to-audio generation, they still face challenges in finding a flexible and precise way to modify target events within an audio track. We present a novel approach, referred to as PPAE, which serves as a general module for diffusion models and enables precise audio editing. The editing is based on the input textual prompt only and is entirely training-free. We exploit the cross-attention maps of diffusion models to facilitate accurate local editing and employ a hierarchical local-global pipeline to ensure a smoother editing process. Experimental results highlight the effectiveness of our method in various editing tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2406_04350
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Prompt-guided Precise Audio Editing with Diffusion Models
Xu, Manjie
Li, Chenxing
zhang, Duzhen
Su, Dan
Liang, Wei
Yu, Dong
Sound
Artificial Intelligence
Machine Learning
Audio and Speech Processing
Audio editing involves the arbitrary manipulation of audio content through precise control. Although text-guided diffusion models have made significant advancements in text-to-audio generation, they still face challenges in finding a flexible and precise way to modify target events within an audio track. We present a novel approach, referred to as PPAE, which serves as a general module for diffusion models and enables precise audio editing. The editing is based on the input textual prompt only and is entirely training-free. We exploit the cross-attention maps of diffusion models to facilitate accurate local editing and employ a hierarchical local-global pipeline to ensure a smoother editing process. Experimental results highlight the effectiveness of our method in various editing tasks.
title Prompt-guided Precise Audio Editing with Diffusion Models
topic Sound
Artificial Intelligence
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2406.04350