FlexIO: Flexible Single- and Multi-Channel Speech Separation and Enhancement
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911229781475328 |
|---|---|
| author | Masuyama, Yoshiki Saijo, Kohei Paissan, Francesco Han, Jiangyu Delcroix, Marc Aihara, Ryo Germain, François G. Wichern, Gordon Roux, Jonathan Le |
| author_facet | Masuyama, Yoshiki Saijo, Kohei Paissan, Francesco Han, Jiangyu Delcroix, Marc Aihara, Ryo Germain, François G. Wichern, Gordon Roux, Jonathan Le |
| contents | Speech separation and enhancement (SSE) has advanced remarkably and achieved promising results in controlled settings, such as a fixed number of speakers and a fixed array configuration. Towards a universal SSE system, single-channel systems have been extended to deal with a variable number of speakers (i.e., outputs). Meanwhile, multi-channel systems accommodating various array configurations (i.e., inputs) have been developed. However, these attempts have been pursued separately. In this paper, we propose a flexible input and output SSE system, named FlexIO. It performs conditional separation using prompt vectors, one per speaker as a condition, allowing separation of an arbitrary number of speakers. Multi-channel mixtures are processed together with the prompt vectors via an array-agnostic channel communication mechanism. Our experiments demonstrate that FlexIO successfully covers diverse conditions with one to five microphones and one to three speakers. We also confirm the robustness of FlexIO on CHiME-4 real data. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_21485 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | FlexIO: Flexible Single- and Multi-Channel Speech Separation and Enhancement Masuyama, Yoshiki Saijo, Kohei Paissan, Francesco Han, Jiangyu Delcroix, Marc Aihara, Ryo Germain, François G. Wichern, Gordon Roux, Jonathan Le Sound Audio and Speech Processing Signal Processing Speech separation and enhancement (SSE) has advanced remarkably and achieved promising results in controlled settings, such as a fixed number of speakers and a fixed array configuration. Towards a universal SSE system, single-channel systems have been extended to deal with a variable number of speakers (i.e., outputs). Meanwhile, multi-channel systems accommodating various array configurations (i.e., inputs) have been developed. However, these attempts have been pursued separately. In this paper, we propose a flexible input and output SSE system, named FlexIO. It performs conditional separation using prompt vectors, one per speaker as a condition, allowing separation of an arbitrary number of speakers. Multi-channel mixtures are processed together with the prompt vectors via an array-agnostic channel communication mechanism. Our experiments demonstrate that FlexIO successfully covers diverse conditions with one to five microphones and one to three speakers. We also confirm the robustness of FlexIO on CHiME-4 real data. |
| title | FlexIO: Flexible Single- and Multi-Channel Speech Separation and Enhancement |
| topic | Sound Audio and Speech Processing Signal Processing |
| url | https://arxiv.org/abs/2510.21485 |