Single-Channel Target Speech Extraction Utilizing Distance and Room Clues
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918026798956544 |
|---|---|
| author | Shi, Runwu Lin, Zirui Yen, Benjamin Wang, Jiang Nihal, Ragib Amin Nakadai, Kazuhiro |
| author_facet | Shi, Runwu Lin, Zirui Yen, Benjamin Wang, Jiang Nihal, Ragib Amin Nakadai, Kazuhiro |
| contents | This paper aims to achieve single-channel target speech extraction (TSE) in enclosures utilizing distance clues and room information. Recent works have verified the feasibility of distance clues for the TSE task, which can imply the sound source's direct-to-reverberation ratio (DRR) and thus can be utilized for speech separation and TSE systems. However, such distance clue is significantly influenced by the room's acoustic characteristics, such as dimension and reverberation time, making it challenging for TSE systems that rely solely on distance clues to generalize across a variety of different rooms. To solve this, we suggest providing room environmental information (room dimensions and reverberation time) for distance-based TSE for better generalization capabilities. Especially, we propose a distance and environment-based TSE model in the time-frequency (TF) domain with learnable distance and room embedding. Results on both simulated and real collected datasets demonstrate its feasibility. Demonstration materials are available at https://runwushi.github.io/distance-room-demo-page/. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_14433 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Single-Channel Target Speech Extraction Utilizing Distance and Room Clues Shi, Runwu Lin, Zirui Yen, Benjamin Wang, Jiang Nihal, Ragib Amin Nakadai, Kazuhiro Audio and Speech Processing Sound This paper aims to achieve single-channel target speech extraction (TSE) in enclosures utilizing distance clues and room information. Recent works have verified the feasibility of distance clues for the TSE task, which can imply the sound source's direct-to-reverberation ratio (DRR) and thus can be utilized for speech separation and TSE systems. However, such distance clue is significantly influenced by the room's acoustic characteristics, such as dimension and reverberation time, making it challenging for TSE systems that rely solely on distance clues to generalize across a variety of different rooms. To solve this, we suggest providing room environmental information (room dimensions and reverberation time) for distance-based TSE for better generalization capabilities. Especially, we propose a distance and environment-based TSE model in the time-frequency (TF) domain with learnable distance and room embedding. Results on both simulated and real collected datasets demonstrate its feasibility. Demonstration materials are available at https://runwushi.github.io/distance-room-demo-page/. |
| title | Single-Channel Target Speech Extraction Utilizing Distance and Room Clues |
| topic | Audio and Speech Processing Sound |
| url | https://arxiv.org/abs/2505.14433 |