Audio Data Preparation and Augmentation using Tensor Flow
Fuente:
Zenodo
Salvato in:
| Autori principali: | , , , , |
|---|---|
| Natura: | Recurso digital |
| Pubblicazione: |
Zenodo
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866901684747239424 |
|---|---|
| author | Goli. Ranga Nadha Rao Chimata Lakshmi Narendra Sirigiri Akash Tadanki Raja Vijayendra Kuraganti Chaitanya |
| author_facet | Goli. Ranga Nadha Rao Chimata Lakshmi Narendra Sirigiri Akash Tadanki Raja Vijayendra Kuraganti Chaitanya |
| contents | Recognition is the basis of speech recognition,and its application is rapidly increasing in keyword spotting, robotics, and smart home surveillance. Because of these advanced applications, improving the accuracy of keyword recognition is crucial. In this paper, we proposed voice conversion (VC)- based augmentation to increase the limited training dataset and a fusion of a convolutional neural network (CNN) and long-short term memory (LSTM) model for robust speaker-independent isolated keyword recognition. Collecting and preparing a sufficient amount of voice data for speaker-independent speech recognition is a tedious and bulky task. In this study, the main intention of voice conversion is to obtain numerous and various human-like keywords' voices that are not identical to the source and target speakers'pronunciation. We examined the performance of the proposed voice conversion augmentation techniques using robust deep neural network algorithms. Original training data, excluding generated voice using other data augmentation and regularization techniques, were considered as the baseline. The results showed that incorporating voice conversion augmentation into the baseline augmentation techniques and applying the CNN-LSTMmodel improved the accuracy of isolated keyword recognition. |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_19335684 |
| institution | Zenodo |
| language | |
| publishDate | 2026 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | Audio Data Preparation and Augmentation using Tensor Flow Goli. Ranga Nadha Rao Chimata Lakshmi Narendra Sirigiri Akash Tadanki Raja Vijayendra Kuraganti Chaitanya Audio Data Processing Audio Augmentation TensorFlow MachineLearning Deep Learning Spectrogram Mel-Spectrogram Feature Extraction Signal Processing Time Shifting Pitch Shifting Time Stretching. Recognition is the basis of speech recognition,and its application is rapidly increasing in keyword spotting, robotics, and smart home surveillance. Because of these advanced applications, improving the accuracy of keyword recognition is crucial. In this paper, we proposed voice conversion (VC)- based augmentation to increase the limited training dataset and a fusion of a convolutional neural network (CNN) and long-short term memory (LSTM) model for robust speaker-independent isolated keyword recognition. Collecting and preparing a sufficient amount of voice data for speaker-independent speech recognition is a tedious and bulky task. In this study, the main intention of voice conversion is to obtain numerous and various human-like keywords' voices that are not identical to the source and target speakers'pronunciation. We examined the performance of the proposed voice conversion augmentation techniques using robust deep neural network algorithms. Original training data, excluding generated voice using other data augmentation and regularization techniques, were considered as the baseline. The results showed that incorporating voice conversion augmentation into the baseline augmentation techniques and applying the CNN-LSTMmodel improved the accuracy of isolated keyword recognition. |
| title | Audio Data Preparation and Augmentation using Tensor Flow |
| topic | Audio Data Processing Audio Augmentation TensorFlow MachineLearning Deep Learning Spectrogram Mel-Spectrogram Feature Extraction Signal Processing Time Shifting Pitch Shifting Time Stretching. |
| url | https://doi.org/10.5281/zenodo.19335684 |