A lightweight dual-stage framework for personalized speech enhancement based on DeepFilterNet2

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Serre, Thomas, Fontaine, Mathieu, Benhaim, Éric, Dutour, Geoffroy, Essid, Slim
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909167290155008
author Serre, Thomas
Fontaine, Mathieu
Benhaim, Éric
Dutour, Geoffroy
Essid, Slim
author_facet Serre, Thomas
Fontaine, Mathieu
Benhaim, Éric
Dutour, Geoffroy
Essid, Slim
contents Isolating the desired speaker's voice amidst multiplespeakers in a noisy acoustic context is a challenging task. Per-sonalized speech enhancement (PSE) endeavours to achievethis by leveraging prior knowledge of the speaker's voice.Recent research efforts have yielded promising PSE mod-els, albeit often accompanied by computationally intensivearchitectures, unsuitable for resource-constrained embeddeddevices. In this paper, we introduce a novel method to per-sonalize a lightweight dual-stage Speech Enhancement (SE)model and implement it within DeepFilterNet2, a SE modelrenowned for its state-of-the-art performance. We seek anoptimal integration of speaker information within the model,exploring different positions for the integration of the speakerembeddings within the dual-stage enhancement architec-ture. We also investigate a tailored training strategy whenadapting DeepFilterNet2 to a PSE task. We show that ourpersonalization method greatly improves the performancesof DeepFilterNet2 while preserving minimal computationaloverhead.
format Preprint
id arxiv_https___arxiv_org_abs_2404_08022
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A lightweight dual-stage framework for personalized speech enhancement based on DeepFilterNet2
Serre, Thomas
Fontaine, Mathieu
Benhaim, Éric
Dutour, Geoffroy
Essid, Slim
Sound
Audio and Speech Processing
Isolating the desired speaker's voice amidst multiplespeakers in a noisy acoustic context is a challenging task. Per-sonalized speech enhancement (PSE) endeavours to achievethis by leveraging prior knowledge of the speaker's voice.Recent research efforts have yielded promising PSE mod-els, albeit often accompanied by computationally intensivearchitectures, unsuitable for resource-constrained embeddeddevices. In this paper, we introduce a novel method to per-sonalize a lightweight dual-stage Speech Enhancement (SE)model and implement it within DeepFilterNet2, a SE modelrenowned for its state-of-the-art performance. We seek anoptimal integration of speaker information within the model,exploring different positions for the integration of the speakerembeddings within the dual-stage enhancement architec-ture. We also investigate a tailored training strategy whenadapting DeepFilterNet2 to a PSE task. We show that ourpersonalization method greatly improves the performancesof DeepFilterNet2 while preserving minimal computationaloverhead.
title A lightweight dual-stage framework for personalized speech enhancement based on DeepFilterNet2
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2404.08022