Optimizing Neural Architectures for Hindi Speech Separation and Enhancement in Noisy Environments

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
1. Verfasser: Ramamoorthy, Arnav
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915456474939392
author Ramamoorthy, Arnav
author_facet Ramamoorthy, Arnav
contents This paper addresses the challenges of Hindi speech separation and enhancement using advanced neural network architectures, with a focus on edge devices. We propose a refined approach leveraging the DEMUCS model to overcome limitations of traditional methods, achieving substantial improvements in speech clarity and intelligibility. The model is fine-tuned with U-Net and LSTM layers, trained on a dataset of 400,000 Hindi speech clips augmented with ESC-50 and MS-SNSD for diverse acoustic environments. Evaluation using PESQ and STOI metrics shows superior performance, particularly under extreme noise conditions. To ensure deployment on resource-constrained devices like TWS earbuds, we explore quantization techniques to reduce computational requirements. This research highlights the effectiveness of customized AI algorithms for speech processing in Indian contexts and suggests future directions for optimizing edge-based architectures.
format Preprint
id arxiv_https___arxiv_org_abs_2508_12009
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Optimizing Neural Architectures for Hindi Speech Separation and Enhancement in Noisy Environments
Ramamoorthy, Arnav
Sound
Machine Learning
This paper addresses the challenges of Hindi speech separation and enhancement using advanced neural network architectures, with a focus on edge devices. We propose a refined approach leveraging the DEMUCS model to overcome limitations of traditional methods, achieving substantial improvements in speech clarity and intelligibility. The model is fine-tuned with U-Net and LSTM layers, trained on a dataset of 400,000 Hindi speech clips augmented with ESC-50 and MS-SNSD for diverse acoustic environments. Evaluation using PESQ and STOI metrics shows superior performance, particularly under extreme noise conditions. To ensure deployment on resource-constrained devices like TWS earbuds, we explore quantization techniques to reduce computational requirements. This research highlights the effectiveness of customized AI algorithms for speech processing in Indian contexts and suggests future directions for optimizing edge-based architectures.
title Optimizing Neural Architectures for Hindi Speech Separation and Enhancement in Noisy Environments
topic Sound
Machine Learning
url https://arxiv.org/abs/2508.12009