A Novel Transfer Learning Approach for Mental Stability Classification from Voice Signal

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Islam, Rafiul, Ahad, Md. Taimur
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915750437978112
author Islam, Rafiul
Ahad, Md. Taimur
author_facet Islam, Rafiul
Ahad, Md. Taimur
contents This study presents a novel transfer learning approach and data augmentation technique for mental stability classification using human voice signals and addresses the challenges associated with limited data availability. Convolutional neural networks (CNNs) have been employed to analyse spectrogram images generated from voice recordings. Three CNN architectures, VGG16, InceptionV3, and DenseNet121, were evaluated across three experimental phases: training on non-augmented data, augmented data, and transfer learning. This proposed transfer learning approach involves pre-training models on the augmented dataset and fine-tuning them on the non-augmented dataset while ensuring strict data separation to prevent data leakage. The results demonstrate significant improvements in classification performance compared to the baseline approach. Among three CNN architectures, DenseNet121 achieved the highest accuracy of 94% and an AUC score of 99% using the proposed transfer learning approach. This finding highlights the effectiveness of combining data augmentation and transfer learning to enhance CNN-based classification of mental stability using voice spectrograms, offering a promising non-invasive tool for mental health diagnostics.
format Preprint
id arxiv_https___arxiv_org_abs_2601_16793
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle A Novel Transfer Learning Approach for Mental Stability Classification from Voice Signal
Islam, Rafiul
Ahad, Md. Taimur
Sound
Neural and Evolutionary Computing
Audio and Speech Processing
This study presents a novel transfer learning approach and data augmentation technique for mental stability classification using human voice signals and addresses the challenges associated with limited data availability. Convolutional neural networks (CNNs) have been employed to analyse spectrogram images generated from voice recordings. Three CNN architectures, VGG16, InceptionV3, and DenseNet121, were evaluated across three experimental phases: training on non-augmented data, augmented data, and transfer learning. This proposed transfer learning approach involves pre-training models on the augmented dataset and fine-tuning them on the non-augmented dataset while ensuring strict data separation to prevent data leakage. The results demonstrate significant improvements in classification performance compared to the baseline approach. Among three CNN architectures, DenseNet121 achieved the highest accuracy of 94% and an AUC score of 99% using the proposed transfer learning approach. This finding highlights the effectiveness of combining data augmentation and transfer learning to enhance CNN-based classification of mental stability using voice spectrograms, offering a promising non-invasive tool for mental health diagnostics.
title A Novel Transfer Learning Approach for Mental Stability Classification from Voice Signal
topic Sound
Neural and Evolutionary Computing
Audio and Speech Processing
url https://arxiv.org/abs/2601.16793