Multi-Channel Multi-Speaker ASR Using Target Speaker's Solo Segment

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Shao, Yiwen, Zhang, Shi-Xiong, Xu, Yong, Yu, Meng, Yu, Dong, Povey, Daniel, Khudanpur, Sanjeev
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911921541742592
author Shao, Yiwen
Zhang, Shi-Xiong
Xu, Yong
Yu, Meng
Yu, Dong
Povey, Daniel
Khudanpur, Sanjeev
author_facet Shao, Yiwen
Zhang, Shi-Xiong
Xu, Yong
Yu, Meng
Yu, Dong
Povey, Daniel
Khudanpur, Sanjeev
contents In the field of multi-channel, multi-speaker Automatic Speech Recognition (ASR), the task of discerning and accurately transcribing a target speaker's speech within background noise remains a formidable challenge. Traditional approaches often rely on microphone array configurations and the information of the target speaker's location or voiceprint. This study introduces the Solo Spatial Feature (Solo-SF), an innovative method that utilizes a target speaker's isolated speech segment to enhance ASR performance, thereby circumventing the need for conventional inputs like microphone array layouts. We explore effective strategies for selecting optimal solo segments, a crucial aspect for Solo-SF's success. Through evaluations conducted on the AliMeeting dataset and AISHELL-1 simulations, Solo-SF demonstrates superior performance over existing techniques, significantly lowering Character Error Rates (CER) in various test conditions. Our findings highlight Solo-SF's potential as an effective solution for addressing the complexities of multi-channel, multi-speaker ASR tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2406_09589
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multi-Channel Multi-Speaker ASR Using Target Speaker's Solo Segment
Shao, Yiwen
Zhang, Shi-Xiong
Xu, Yong
Yu, Meng
Yu, Dong
Povey, Daniel
Khudanpur, Sanjeev
Audio and Speech Processing
In the field of multi-channel, multi-speaker Automatic Speech Recognition (ASR), the task of discerning and accurately transcribing a target speaker's speech within background noise remains a formidable challenge. Traditional approaches often rely on microphone array configurations and the information of the target speaker's location or voiceprint. This study introduces the Solo Spatial Feature (Solo-SF), an innovative method that utilizes a target speaker's isolated speech segment to enhance ASR performance, thereby circumventing the need for conventional inputs like microphone array layouts. We explore effective strategies for selecting optimal solo segments, a crucial aspect for Solo-SF's success. Through evaluations conducted on the AliMeeting dataset and AISHELL-1 simulations, Solo-SF demonstrates superior performance over existing techniques, significantly lowering Character Error Rates (CER) in various test conditions. Our findings highlight Solo-SF's potential as an effective solution for addressing the complexities of multi-channel, multi-speaker ASR tasks.
title Multi-Channel Multi-Speaker ASR Using Target Speaker's Solo Segment
topic Audio and Speech Processing
url https://arxiv.org/abs/2406.09589