Voice recognition by deep transfer learning and vision transformers to secure voice authentication

Fuente: Zenodo
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Nayem Uddin Prince, Abdullah Al Masum, Salman Mohammad Abdullah, Touhid Bhuiyan
Format: Recurso digital
Langue:anglais
Publié: Zenodo 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866901096471986176
author Nayem Uddin Prince
Abdullah Al Masum
Salman Mohammad Abdullah
Touhid Bhuiyan
author_facet Nayem Uddin Prince
Abdullah Al Masum
Salman Mohammad Abdullah
Touhid Bhuiyan
contents <p>Speech recognition is crucial for ensuring the security of personal devices and financial transactions. Attaining high accuracy and robustness in voice authentication is challenging due to the presence of voice and environmental variability. Recent advancements in the field of deep learning, particularly in transfer learning and visual transformers, have the potential to enhance voice recognition systems. This study employs advanced deep transfer learning techniques, including Vision Transformers (ViT), VGG16, and a customized Convolutional Neural Network (CNN), to enhance the accuracy and security of speech authentication. The objective is to evaluate and contrast various solutions' voice recognition and authentication accuracy. The experiment included 3000 voice samples, with an equal distribution of 1500 samples from male participants and 1500 from female participants. The dataset was used to train Vision Transformers, VGG16 with transfer learning, and a custom CNN. The models were assessed based on their accuracy in identifying and authenticating voice samples. The VGG16 model achieved the highest level of accuracy in speech recognition, with a precision rate of 95%. The Vision Transformer and custom CNN exhibited satisfactory performance. However, VGG16 demonstrated higher accuracy. The most accurate voice authentication model studied is the VGG16 model based on transfer learning. This study suggests that the security and reliability of voice recognition systems can be enhanced through the use of deep learning techniques.</p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_14945303
institution Zenodo
language eng
publishDate 2024
publisher Zenodo
record_format zenodo
spellingShingle Voice recognition by deep transfer learning and vision transformers to secure voice authentication
Nayem Uddin Prince
Abdullah Al Masum
Salman Mohammad Abdullah
Touhid Bhuiyan
Voice recognition
VGG16
CustomCNN
Vit
honey trap
webform
Cybercrime
Vision Transform
MFCCs
<p>Speech recognition is crucial for ensuring the security of personal devices and financial transactions. Attaining high accuracy and robustness in voice authentication is challenging due to the presence of voice and environmental variability. Recent advancements in the field of deep learning, particularly in transfer learning and visual transformers, have the potential to enhance voice recognition systems. This study employs advanced deep transfer learning techniques, including Vision Transformers (ViT), VGG16, and a customized Convolutional Neural Network (CNN), to enhance the accuracy and security of speech authentication. The objective is to evaluate and contrast various solutions' voice recognition and authentication accuracy. The experiment included 3000 voice samples, with an equal distribution of 1500 samples from male participants and 1500 from female participants. The dataset was used to train Vision Transformers, VGG16 with transfer learning, and a custom CNN. The models were assessed based on their accuracy in identifying and authenticating voice samples. The VGG16 model achieved the highest level of accuracy in speech recognition, with a precision rate of 95%. The Vision Transformer and custom CNN exhibited satisfactory performance. However, VGG16 demonstrated higher accuracy. The most accurate voice authentication model studied is the VGG16 model based on transfer learning. This study suggests that the security and reliability of voice recognition systems can be enhanced through the use of deep learning techniques.</p>
title Voice recognition by deep transfer learning and vision transformers to secure voice authentication
topic Voice recognition
VGG16
CustomCNN
Vit
honey trap
webform
Cybercrime
Vision Transform
MFCCs
url https://doi.org/10.5281/zenodo.14945303