Improving Pretrained YAMNet for Enhanced Speech Command Detection via Transfer Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lachenani, Sidahmed, Kheddar, Hamza, Ouldzmirli, Mohamed
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916709037768704
author Lachenani, Sidahmed
Kheddar, Hamza
Ouldzmirli, Mohamed
author_facet Lachenani, Sidahmed
Kheddar, Hamza
Ouldzmirli, Mohamed
contents This work addresses the need for enhanced accuracy and efficiency in speech command recognition systems, a critical component for improving user interaction in various smart applications. Leveraging the robust pretrained YAMNet model and transfer learning, this study develops a method that significantly improves speech command recognition. We adapt and train a YAMNet deep learning model to effectively detect and interpret speech commands from audio signals. Using the extensively annotated Speech Commands dataset (speech_commands_v0.01), our approach demonstrates the practical application of transfer learning to accurately recognize a predefined set of speech commands. The dataset is meticulously augmented, and features are strategically extracted to boost model performance. As a result, the final model achieved a recognition accuracy of 95.28%, underscoring the impact of advanced machine learning techniques on speech command recognition. This achievement marks substantial progress in audio processing technologies and establishes a new benchmark for future research in the field.
format Preprint
id arxiv_https___arxiv_org_abs_2504_19030
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Improving Pretrained YAMNet for Enhanced Speech Command Detection via Transfer Learning
Lachenani, Sidahmed
Kheddar, Hamza
Ouldzmirli, Mohamed
Sound
Artificial Intelligence
Audio and Speech Processing
This work addresses the need for enhanced accuracy and efficiency in speech command recognition systems, a critical component for improving user interaction in various smart applications. Leveraging the robust pretrained YAMNet model and transfer learning, this study develops a method that significantly improves speech command recognition. We adapt and train a YAMNet deep learning model to effectively detect and interpret speech commands from audio signals. Using the extensively annotated Speech Commands dataset (speech_commands_v0.01), our approach demonstrates the practical application of transfer learning to accurately recognize a predefined set of speech commands. The dataset is meticulously augmented, and features are strategically extracted to boost model performance. As a result, the final model achieved a recognition accuracy of 95.28%, underscoring the impact of advanced machine learning techniques on speech command recognition. This achievement marks substantial progress in audio processing technologies and establishes a new benchmark for future research in the field.
title Improving Pretrained YAMNet for Enhanced Speech Command Detection via Transfer Learning
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2504.19030