Can Synthetic Audio From Generative Foundation Models Assist Audio Recognition and Speech Modeling?

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Feng, Tiantian, Dimitriadis, Dimitrios, Narayanan, Shrikanth
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913485380648960
author Feng, Tiantian
Dimitriadis, Dimitrios
Narayanan, Shrikanth
author_facet Feng, Tiantian
Dimitriadis, Dimitrios
Narayanan, Shrikanth
contents Recent advances in foundation models have enabled audio-generative models that produce high-fidelity sounds associated with music, events, and human actions. Despite the success achieved in modern audio-generative models, the conventional approach to assessing the quality of the audio generation relies heavily on distance metrics like Frechet Audio Distance. In contrast, we aim to evaluate the quality of audio generation by examining the effectiveness of using them as training data. Specifically, we conduct studies to explore the use of synthetic audio for audio recognition. Moreover, we investigate whether synthetic audio can serve as a resource for data augmentation in speech-related modeling. Our comprehensive experiments demonstrate the potential of using synthetic audio for audio recognition and speech-related modeling. Our code is available at https://github.com/usc-sail/SynthAudio.
format Preprint
id arxiv_https___arxiv_org_abs_2406_08800
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Can Synthetic Audio From Generative Foundation Models Assist Audio Recognition and Speech Modeling?
Feng, Tiantian
Dimitriadis, Dimitrios
Narayanan, Shrikanth
Sound
Machine Learning
Audio and Speech Processing
Recent advances in foundation models have enabled audio-generative models that produce high-fidelity sounds associated with music, events, and human actions. Despite the success achieved in modern audio-generative models, the conventional approach to assessing the quality of the audio generation relies heavily on distance metrics like Frechet Audio Distance. In contrast, we aim to evaluate the quality of audio generation by examining the effectiveness of using them as training data. Specifically, we conduct studies to explore the use of synthetic audio for audio recognition. Moreover, we investigate whether synthetic audio can serve as a resource for data augmentation in speech-related modeling. Our comprehensive experiments demonstrate the potential of using synthetic audio for audio recognition and speech-related modeling. Our code is available at https://github.com/usc-sail/SynthAudio.
title Can Synthetic Audio From Generative Foundation Models Assist Audio Recognition and Speech Modeling?
topic Sound
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2406.08800