aTENNuate: Optimized Real-time Speech Enhancement with Deep SSMs on Raw Audio

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Pei, Yan Ru, Shrivastava, Ritik, Sidharth, FNU
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913892976820224
author Pei, Yan Ru
Shrivastava, Ritik
Sidharth, FNU
author_facet Pei, Yan Ru
Shrivastava, Ritik
Sidharth, FNU
contents We present aTENNuate, a simple deep state-space autoencoder configured for efficient online raw speech enhancement in an end-to-end fashion. The network's performance is primarily evaluated on raw speech denoising, with additional assessments on tasks such as super-resolution and de-quantization. We benchmark aTENNuate on the VoiceBank + DEMAND and the Microsoft DNS1 synthetic test sets. The network outperforms previous real-time denoising models in terms of PESQ score, parameter count, MACs, and latency. Even as a raw waveform processing model, the model maintains high fidelity to the clean signal with minimal audible artifacts. In addition, the model remains performant even when the noisy input is compressed down to 4000Hz and 4 bits, suggesting general speech enhancement capabilities in low-resource environments. Try it out by pip install attenuate
format Preprint
id arxiv_https___arxiv_org_abs_2409_03377
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle aTENNuate: Optimized Real-time Speech Enhancement with Deep SSMs on Raw Audio
Pei, Yan Ru
Shrivastava, Ritik
Sidharth, FNU
Sound
Artificial Intelligence
Machine Learning
Audio and Speech Processing
We present aTENNuate, a simple deep state-space autoencoder configured for efficient online raw speech enhancement in an end-to-end fashion. The network's performance is primarily evaluated on raw speech denoising, with additional assessments on tasks such as super-resolution and de-quantization. We benchmark aTENNuate on the VoiceBank + DEMAND and the Microsoft DNS1 synthetic test sets. The network outperforms previous real-time denoising models in terms of PESQ score, parameter count, MACs, and latency. Even as a raw waveform processing model, the model maintains high fidelity to the clean signal with minimal audible artifacts. In addition, the model remains performant even when the noisy input is compressed down to 4000Hz and 4 bits, suggesting general speech enhancement capabilities in low-resource environments. Try it out by pip install attenuate
title aTENNuate: Optimized Real-time Speech Enhancement with Deep SSMs on Raw Audio
topic Sound
Artificial Intelligence
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2409.03377