Spectron: Target Speaker Extraction using Conditional Transformer with Adversarial Refinement

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Bandyopadhyay, Tathagata
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912012751077376
author Bandyopadhyay, Tathagata
author_facet Bandyopadhyay, Tathagata
contents Recently, attention-based transformers have become a de facto standard in many deep learning applications including natural language processing, computer vision, signal processing, etc.. In this paper, we propose a transformer-based end-to-end model to extract a target speaker's speech from a monaural multi-speaker mixed audio signal. Unlike existing speaker extraction methods, we introduce two additional objectives to impose speaker embedding consistency and waveform encoder invertibility and jointly train both speaker encoder and speech separator to better capture the speaker conditional embedding. Furthermore, we leverage a multi-scale discriminator to refine the perceptual quality of the extracted speech. Our experiments show that the use of a dual path transformer in the separator backbone along with proposed training paradigm improves the CNN baseline by $3.12$ dB points. Finally, we compare our approach with recent state-of-the-arts and show that our model outperforms existing methods by $4.1$ dB points on an average without creating additional data dependency.
format Preprint
id arxiv_https___arxiv_org_abs_2409_01352
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Spectron: Target Speaker Extraction using Conditional Transformer with Adversarial Refinement
Bandyopadhyay, Tathagata
Sound
Machine Learning
Multimedia
Audio and Speech Processing
Recently, attention-based transformers have become a de facto standard in many deep learning applications including natural language processing, computer vision, signal processing, etc.. In this paper, we propose a transformer-based end-to-end model to extract a target speaker's speech from a monaural multi-speaker mixed audio signal. Unlike existing speaker extraction methods, we introduce two additional objectives to impose speaker embedding consistency and waveform encoder invertibility and jointly train both speaker encoder and speech separator to better capture the speaker conditional embedding. Furthermore, we leverage a multi-scale discriminator to refine the perceptual quality of the extracted speech. Our experiments show that the use of a dual path transformer in the separator backbone along with proposed training paradigm improves the CNN baseline by $3.12$ dB points. Finally, we compare our approach with recent state-of-the-arts and show that our model outperforms existing methods by $4.1$ dB points on an average without creating additional data dependency.
title Spectron: Target Speaker Extraction using Conditional Transformer with Adversarial Refinement
topic Sound
Machine Learning
Multimedia
Audio and Speech Processing
url https://arxiv.org/abs/2409.01352