Audio Data Preparation and Augmentation using Tensor Flow

Fuente: Zenodo
Salvato in:
Dettagli Bibliografici
Autori principali: Goli. Ranga Nadha Rao, Chimata Lakshmi Narendra, Sirigiri Akash, Tadanki Raja Vijayendra, Kuraganti Chaitanya
Natura: Recurso digital
Pubblicazione: Zenodo 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866901684747239424
author Goli. Ranga Nadha Rao
Chimata Lakshmi Narendra
Sirigiri Akash
Tadanki Raja Vijayendra
Kuraganti Chaitanya
author_facet Goli. Ranga Nadha Rao
Chimata Lakshmi Narendra
Sirigiri Akash
Tadanki Raja Vijayendra
Kuraganti Chaitanya
contents Recognition is the basis of speech recognition,and its application is rapidly increasing in keyword spotting, robotics, and smart home surveillance. Because of these advanced applications, improving the accuracy of keyword recognition is crucial. In this paper, we proposed voice conversion (VC)- based augmentation to increase the limited training dataset and a fusion of a convolutional neural network (CNN) and long-short term memory (LSTM) model for robust speaker-independent isolated keyword recognition. Collecting and preparing a sufficient amount of voice data for speaker-independent speech recognition is a tedious and bulky task. In this study, the main intention of voice conversion is to obtain numerous and various human-like keywords' voices that are not identical to the source and target speakers'pronunciation. We examined the performance of the proposed voice conversion augmentation techniques using robust deep neural network algorithms. Original training data, excluding generated voice using other data augmentation and regularization techniques, were considered as the baseline. The results showed that incorporating voice conversion augmentation into the baseline augmentation techniques and applying the CNN-LSTMmodel improved the accuracy of isolated keyword recognition.
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_19335684
institution Zenodo
language
publishDate 2026
publisher Zenodo
record_format zenodo
spellingShingle Audio Data Preparation and Augmentation using Tensor Flow
Goli. Ranga Nadha Rao
Chimata Lakshmi Narendra
Sirigiri Akash
Tadanki Raja Vijayendra
Kuraganti Chaitanya
Audio Data Processing
Audio Augmentation
TensorFlow
MachineLearning
Deep Learning
Spectrogram
Mel-Spectrogram
Feature Extraction
Signal Processing
Time Shifting
Pitch Shifting
Time Stretching.
Recognition is the basis of speech recognition,and its application is rapidly increasing in keyword spotting, robotics, and smart home surveillance. Because of these advanced applications, improving the accuracy of keyword recognition is crucial. In this paper, we proposed voice conversion (VC)- based augmentation to increase the limited training dataset and a fusion of a convolutional neural network (CNN) and long-short term memory (LSTM) model for robust speaker-independent isolated keyword recognition. Collecting and preparing a sufficient amount of voice data for speaker-independent speech recognition is a tedious and bulky task. In this study, the main intention of voice conversion is to obtain numerous and various human-like keywords' voices that are not identical to the source and target speakers'pronunciation. We examined the performance of the proposed voice conversion augmentation techniques using robust deep neural network algorithms. Original training data, excluding generated voice using other data augmentation and regularization techniques, were considered as the baseline. The results showed that incorporating voice conversion augmentation into the baseline augmentation techniques and applying the CNN-LSTMmodel improved the accuracy of isolated keyword recognition.
title Audio Data Preparation and Augmentation using Tensor Flow
topic Audio Data Processing
Audio Augmentation
TensorFlow
MachineLearning
Deep Learning
Spectrogram
Mel-Spectrogram
Feature Extraction
Signal Processing
Time Shifting
Pitch Shifting
Time Stretching.
url https://doi.org/10.5281/zenodo.19335684