Target Speech Extraction with Pre-trained AV-HuBERT and Mask-And-Recover Strategy

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Wenxuan, Chen, Xueyuan, Wu, Xixin, Li, Haizhou, Meng, Helen
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917621460369408
author Wu, Wenxuan
Chen, Xueyuan
Wu, Xixin
Li, Haizhou
Meng, Helen
author_facet Wu, Wenxuan
Chen, Xueyuan
Wu, Xixin
Li, Haizhou
Meng, Helen
contents Audio-visual target speech extraction (AV-TSE) is one of the enabling technologies in robotics and many audio-visual applications. One of the challenges of AV-TSE is how to effectively utilize audio-visual synchronization information in the process. AV-HuBERT can be a useful pre-trained model for lip-reading, which has not been adopted by AV-TSE. In this paper, we would like to explore the way to integrate a pre-trained AV-HuBERT into our AV-TSE system. We have good reasons to expect an improved performance. To benefit from the inter and intra-modality correlations, we also propose a novel Mask-And-Recover (MAR) strategy for self-supervised learning. The experimental results on the VoxCeleb2 dataset show that our proposed model outperforms the baselines both in terms of subjective and objective metrics, suggesting that the pre-trained AV-HuBERT model provides more informative visual cues for target speech extraction. Furthermore, through a comparative study, we confirm that the proposed Mask-And-Recover strategy is significantly effective.
format Preprint
id arxiv_https___arxiv_org_abs_2403_16078
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Target Speech Extraction with Pre-trained AV-HuBERT and Mask-And-Recover Strategy
Wu, Wenxuan
Chen, Xueyuan
Wu, Xixin
Li, Haizhou
Meng, Helen
Sound
Audio and Speech Processing
Audio-visual target speech extraction (AV-TSE) is one of the enabling technologies in robotics and many audio-visual applications. One of the challenges of AV-TSE is how to effectively utilize audio-visual synchronization information in the process. AV-HuBERT can be a useful pre-trained model for lip-reading, which has not been adopted by AV-TSE. In this paper, we would like to explore the way to integrate a pre-trained AV-HuBERT into our AV-TSE system. We have good reasons to expect an improved performance. To benefit from the inter and intra-modality correlations, we also propose a novel Mask-And-Recover (MAR) strategy for self-supervised learning. The experimental results on the VoxCeleb2 dataset show that our proposed model outperforms the baselines both in terms of subjective and objective metrics, suggesting that the pre-trained AV-HuBERT model provides more informative visual cues for target speech extraction. Furthermore, through a comparative study, we confirm that the proposed Mask-And-Recover strategy is significantly effective.
title Target Speech Extraction with Pre-trained AV-HuBERT and Mask-And-Recover Strategy
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2403.16078