End-to-end multi-channel speaker extraction and binaural speech synthesis

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Chi, Cheng, Li, Xiaoyu, Ke, Yuxuan, Ni, Qunping, Ge, Yao, Li, Xiaodong, Zheng, Chengshi
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912476691431424
author Chi, Cheng
Li, Xiaoyu
Ke, Yuxuan
Ni, Qunping
Ge, Yao
Li, Xiaodong
Zheng, Chengshi
author_facet Chi, Cheng
Li, Xiaoyu
Ke, Yuxuan
Ni, Qunping
Ge, Yao
Li, Xiaodong
Zheng, Chengshi
contents Speech clarity and spatial audio immersion are the two most critical factors in enhancing remote conferencing experiences. Existing methods are often limited: either due to the lack of spatial information when using only one microphone, or because their performance is highly dependent on the accuracy of direction-of-arrival estimation when using microphone array. To overcome this issue, we introduce an end-to-end deep learning framework that has the capacity of mapping multi-channel noisy and reverberant signals to clean and spatialized binaural speech directly. This framework unifies source extraction, noise suppression, and binaural rendering into one network. In this framework, a novel magnitude-weighted interaural level difference loss function is proposed that aims to improve the accuracy of spatial rendering. Extensive evaluations show that our method outperforms established baselines in terms of both speech quality and spatial fidelity.
format Preprint
id arxiv_https___arxiv_org_abs_2410_05739
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle End-to-end multi-channel speaker extraction and binaural speech synthesis
Chi, Cheng
Li, Xiaoyu
Ke, Yuxuan
Ni, Qunping
Ge, Yao
Li, Xiaodong
Zheng, Chengshi
Sound
Artificial Intelligence
Audio and Speech Processing
Speech clarity and spatial audio immersion are the two most critical factors in enhancing remote conferencing experiences. Existing methods are often limited: either due to the lack of spatial information when using only one microphone, or because their performance is highly dependent on the accuracy of direction-of-arrival estimation when using microphone array. To overcome this issue, we introduce an end-to-end deep learning framework that has the capacity of mapping multi-channel noisy and reverberant signals to clean and spatialized binaural speech directly. This framework unifies source extraction, noise suppression, and binaural rendering into one network. In this framework, a novel magnitude-weighted interaural level difference loss function is proposed that aims to improve the accuracy of spatial rendering. Extensive evaluations show that our method outperforms established baselines in terms of both speech quality and spatial fidelity.
title End-to-end multi-channel speaker extraction and binaural speech synthesis
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2410.05739