Single-Channel Target Speech Extraction Utilizing Distance and Room Clues

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shi, Runwu, Lin, Zirui, Yen, Benjamin, Wang, Jiang, Nihal, Ragib Amin, Nakadai, Kazuhiro
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918026798956544
author Shi, Runwu
Lin, Zirui
Yen, Benjamin
Wang, Jiang
Nihal, Ragib Amin
Nakadai, Kazuhiro
author_facet Shi, Runwu
Lin, Zirui
Yen, Benjamin
Wang, Jiang
Nihal, Ragib Amin
Nakadai, Kazuhiro
contents This paper aims to achieve single-channel target speech extraction (TSE) in enclosures utilizing distance clues and room information. Recent works have verified the feasibility of distance clues for the TSE task, which can imply the sound source's direct-to-reverberation ratio (DRR) and thus can be utilized for speech separation and TSE systems. However, such distance clue is significantly influenced by the room's acoustic characteristics, such as dimension and reverberation time, making it challenging for TSE systems that rely solely on distance clues to generalize across a variety of different rooms. To solve this, we suggest providing room environmental information (room dimensions and reverberation time) for distance-based TSE for better generalization capabilities. Especially, we propose a distance and environment-based TSE model in the time-frequency (TF) domain with learnable distance and room embedding. Results on both simulated and real collected datasets demonstrate its feasibility. Demonstration materials are available at https://runwushi.github.io/distance-room-demo-page/.
format Preprint
id arxiv_https___arxiv_org_abs_2505_14433
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Single-Channel Target Speech Extraction Utilizing Distance and Room Clues
Shi, Runwu
Lin, Zirui
Yen, Benjamin
Wang, Jiang
Nihal, Ragib Amin
Nakadai, Kazuhiro
Audio and Speech Processing
Sound
This paper aims to achieve single-channel target speech extraction (TSE) in enclosures utilizing distance clues and room information. Recent works have verified the feasibility of distance clues for the TSE task, which can imply the sound source's direct-to-reverberation ratio (DRR) and thus can be utilized for speech separation and TSE systems. However, such distance clue is significantly influenced by the room's acoustic characteristics, such as dimension and reverberation time, making it challenging for TSE systems that rely solely on distance clues to generalize across a variety of different rooms. To solve this, we suggest providing room environmental information (room dimensions and reverberation time) for distance-based TSE for better generalization capabilities. Especially, we propose a distance and environment-based TSE model in the time-frequency (TF) domain with learnable distance and room embedding. Results on both simulated and real collected datasets demonstrate its feasibility. Demonstration materials are available at https://runwushi.github.io/distance-room-demo-page/.
title Single-Channel Target Speech Extraction Utilizing Distance and Room Clues
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2505.14433