FlexIO: Flexible Single- and Multi-Channel Speech Separation and Enhancement

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Masuyama, Yoshiki, Saijo, Kohei, Paissan, Francesco, Han, Jiangyu, Delcroix, Marc, Aihara, Ryo, Germain, François G., Wichern, Gordon, Roux, Jonathan Le
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911229781475328
author Masuyama, Yoshiki
Saijo, Kohei
Paissan, Francesco
Han, Jiangyu
Delcroix, Marc
Aihara, Ryo
Germain, François G.
Wichern, Gordon
Roux, Jonathan Le
author_facet Masuyama, Yoshiki
Saijo, Kohei
Paissan, Francesco
Han, Jiangyu
Delcroix, Marc
Aihara, Ryo
Germain, François G.
Wichern, Gordon
Roux, Jonathan Le
contents Speech separation and enhancement (SSE) has advanced remarkably and achieved promising results in controlled settings, such as a fixed number of speakers and a fixed array configuration. Towards a universal SSE system, single-channel systems have been extended to deal with a variable number of speakers (i.e., outputs). Meanwhile, multi-channel systems accommodating various array configurations (i.e., inputs) have been developed. However, these attempts have been pursued separately. In this paper, we propose a flexible input and output SSE system, named FlexIO. It performs conditional separation using prompt vectors, one per speaker as a condition, allowing separation of an arbitrary number of speakers. Multi-channel mixtures are processed together with the prompt vectors via an array-agnostic channel communication mechanism. Our experiments demonstrate that FlexIO successfully covers diverse conditions with one to five microphones and one to three speakers. We also confirm the robustness of FlexIO on CHiME-4 real data.
format Preprint
id arxiv_https___arxiv_org_abs_2510_21485
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FlexIO: Flexible Single- and Multi-Channel Speech Separation and Enhancement
Masuyama, Yoshiki
Saijo, Kohei
Paissan, Francesco
Han, Jiangyu
Delcroix, Marc
Aihara, Ryo
Germain, François G.
Wichern, Gordon
Roux, Jonathan Le
Sound
Audio and Speech Processing
Signal Processing
Speech separation and enhancement (SSE) has advanced remarkably and achieved promising results in controlled settings, such as a fixed number of speakers and a fixed array configuration. Towards a universal SSE system, single-channel systems have been extended to deal with a variable number of speakers (i.e., outputs). Meanwhile, multi-channel systems accommodating various array configurations (i.e., inputs) have been developed. However, these attempts have been pursued separately. In this paper, we propose a flexible input and output SSE system, named FlexIO. It performs conditional separation using prompt vectors, one per speaker as a condition, allowing separation of an arbitrary number of speakers. Multi-channel mixtures are processed together with the prompt vectors via an array-agnostic channel communication mechanism. Our experiments demonstrate that FlexIO successfully covers diverse conditions with one to five microphones and one to three speakers. We also confirm the robustness of FlexIO on CHiME-4 real data.
title FlexIO: Flexible Single- and Multi-Channel Speech Separation and Enhancement
topic Sound
Audio and Speech Processing
Signal Processing
url https://arxiv.org/abs/2510.21485