Audio-Visual Speech Separation via Bottleneck Iterative Network

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Sidong, Shankar, Shiv, Nguyen, Trang, Fanelli, Andrea, Fiterau, Madalina
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912473605472256
author Zhang, Sidong
Shankar, Shiv
Nguyen, Trang
Fanelli, Andrea
Fiterau, Madalina
author_facet Zhang, Sidong
Shankar, Shiv
Nguyen, Trang
Fanelli, Andrea
Fiterau, Madalina
contents Integration of information from non-auditory cues can significantly improve the performance of speech-separation models. Often such models use deep modality-specific networks to obtain unimodal features, and risk being too costly or lightweight but lacking capacity. In this work, we present an iterative representation refinement approach called Bottleneck Iterative Network (BIN), a technique that repeatedly progresses through a lightweight fusion block, while bottlenecking fusion representations by fusion tokens. This helps improve the capacity of the model, while avoiding major increase in model size and balancing between the model performance and training cost. We test BIN on challenging noisy audio-visual speech separation tasks, and show that our approach consistently outperforms state-of-the-art benchmark models with respect to SI-SDRi on NTCD-TIMIT and LRS3+WHAM! datasets, while simultaneously achieving a reduction of more than 50% in training and GPU inference time across nearly all settings.
format Preprint
id arxiv_https___arxiv_org_abs_2507_07270
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Audio-Visual Speech Separation via Bottleneck Iterative Network
Zhang, Sidong
Shankar, Shiv
Nguyen, Trang
Fanelli, Andrea
Fiterau, Madalina
Sound
Multimedia
Audio and Speech Processing
Integration of information from non-auditory cues can significantly improve the performance of speech-separation models. Often such models use deep modality-specific networks to obtain unimodal features, and risk being too costly or lightweight but lacking capacity. In this work, we present an iterative representation refinement approach called Bottleneck Iterative Network (BIN), a technique that repeatedly progresses through a lightweight fusion block, while bottlenecking fusion representations by fusion tokens. This helps improve the capacity of the model, while avoiding major increase in model size and balancing between the model performance and training cost. We test BIN on challenging noisy audio-visual speech separation tasks, and show that our approach consistently outperforms state-of-the-art benchmark models with respect to SI-SDRi on NTCD-TIMIT and LRS3+WHAM! datasets, while simultaneously achieving a reduction of more than 50% in training and GPU inference time across nearly all settings.
title Audio-Visual Speech Separation via Bottleneck Iterative Network
topic Sound
Multimedia
Audio and Speech Processing
url https://arxiv.org/abs/2507.07270