SRC-gAudio: Sampling-Rate-Controlled Audio Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Chenxing, Xu, Manjie, Yu, Dong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910641763123200
author Li, Chenxing
Xu, Manjie
Yu, Dong
author_facet Li, Chenxing
Xu, Manjie
Yu, Dong
contents We introduce SRC-gAudio, a novel audio generation model designed to facilitate text-to-audio generation across a wide range of sampling rates within a single model architecture. SRC-gAudio incorporates the sampling rate as part of the generation condition to guide the diffusion-based audio generation process. Our model enables the generation of audio at multiple sampling rates with a single unified model. Furthermore, we explore the potential benefits of large-scale, low-sampling-rate data in enhancing the generation quality of high-sampling-rate audio. Through extensive experiments, we demonstrate that SRC-gAudio effectively generates audio under controlled sampling rates. Additionally, our results indicate that pre-training on low-sampling-rate data can lead to significant improvements in audio quality across various metrics.
format Preprint
id arxiv_https___arxiv_org_abs_2410_06544
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SRC-gAudio: Sampling-Rate-Controlled Audio Generation
Li, Chenxing
Xu, Manjie
Yu, Dong
Sound
Audio and Speech Processing
We introduce SRC-gAudio, a novel audio generation model designed to facilitate text-to-audio generation across a wide range of sampling rates within a single model architecture. SRC-gAudio incorporates the sampling rate as part of the generation condition to guide the diffusion-based audio generation process. Our model enables the generation of audio at multiple sampling rates with a single unified model. Furthermore, we explore the potential benefits of large-scale, low-sampling-rate data in enhancing the generation quality of high-sampling-rate audio. Through extensive experiments, we demonstrate that SRC-gAudio effectively generates audio under controlled sampling rates. Additionally, our results indicate that pre-training on low-sampling-rate data can lead to significant improvements in audio quality across various metrics.
title SRC-gAudio: Sampling-Rate-Controlled Audio Generation
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2410.06544