Online Audio-Visual Autoregressive Speaker Extraction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pan, Zexu, Wang, Wupeng, Zhao, Shengkui, Zhang, Chong, Zhou, Kun, Ma, Yukun, Ma, Bin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915318061858816
author Pan, Zexu
Wang, Wupeng
Zhao, Shengkui
Zhang, Chong
Zhou, Kun
Ma, Yukun
Ma, Bin
author_facet Pan, Zexu
Wang, Wupeng
Zhao, Shengkui
Zhang, Chong
Zhou, Kun
Ma, Yukun
Ma, Bin
contents This paper proposes a novel online audio-visual speaker extraction model. In the streaming regime, most studies optimize the audio network only, leaving the visual frontend less explored. We first propose a lightweight visual frontend based on depth-wise separable convolution. Then, we propose a lightweight autoregressive acoustic encoder to serve as the second cue, to actively explore the information in the separated speech signal from past steps. Scenario-wise, for the first time, we study how the algorithm performs when there is a change in focus of attention, i.e., the target speaker. Experimental results on LRS3 datasets show that our visual frontend performs comparably to the previous state-of-the-art on both SkiM and ConvTasNet audio backbones with only 0.1 million network parameters and 2.1 MACs per second of processing. The autoregressive acoustic encoder provides an additional 0.9 dB gain in terms of SI-SNRi, and its momentum is robust against the change in attention.
format Preprint
id arxiv_https___arxiv_org_abs_2506_01270
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Online Audio-Visual Autoregressive Speaker Extraction
Pan, Zexu
Wang, Wupeng
Zhao, Shengkui
Zhang, Chong
Zhou, Kun
Ma, Yukun
Ma, Bin
Audio and Speech Processing
Sound
This paper proposes a novel online audio-visual speaker extraction model. In the streaming regime, most studies optimize the audio network only, leaving the visual frontend less explored. We first propose a lightweight visual frontend based on depth-wise separable convolution. Then, we propose a lightweight autoregressive acoustic encoder to serve as the second cue, to actively explore the information in the separated speech signal from past steps. Scenario-wise, for the first time, we study how the algorithm performs when there is a change in focus of attention, i.e., the target speaker. Experimental results on LRS3 datasets show that our visual frontend performs comparably to the previous state-of-the-art on both SkiM and ConvTasNet audio backbones with only 0.1 million network parameters and 2.1 MACs per second of processing. The autoregressive acoustic encoder provides an additional 0.9 dB gain in terms of SI-SNRi, and its momentum is robust against the change in attention.
title Online Audio-Visual Autoregressive Speaker Extraction
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2506.01270