ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nguyen, Thai-Binh, Van Nguyen, Thi, Do, Quoc Truong, Luong, Chi Mai
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913877463138304
author Nguyen, Thai-Binh
Van Nguyen, Thi
Do, Quoc Truong
Luong, Chi Mai
author_facet Nguyen, Thai-Binh
Van Nguyen, Thi
Do, Quoc Truong
Luong, Chi Mai
contents Audio-Visual Speech Recognition (AVSR) has gained significant attention recently due to its robustness against noise, which often challenges conventional speech recognition systems that rely solely on audio features. Despite this advantage, AVSR models remain limited by the scarcity of extensive datasets, especially for most languages beyond English. Automated data collection offers a promising solution. This work presents a practical approach to generate AVSR datasets from raw video, refining existing techniques for improved efficiency and accessibility. We demonstrate its broad applicability by developing a baseline AVSR model for Vietnamese. Experiments show the automatically collected dataset enables a strong baseline, achieving competitive performance with robust ASR in clean conditions and significantly outperforming them in noisy environments like cocktail parties. This efficient method provides a pathway to expand AVSR to more languages, particularly under-resourced ones.
format Preprint
id arxiv_https___arxiv_org_abs_2506_04635
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition
Nguyen, Thai-Binh
Van Nguyen, Thi
Do, Quoc Truong
Luong, Chi Mai
Computation and Language
Computer Vision and Pattern Recognition
Audio-Visual Speech Recognition (AVSR) has gained significant attention recently due to its robustness against noise, which often challenges conventional speech recognition systems that rely solely on audio features. Despite this advantage, AVSR models remain limited by the scarcity of extensive datasets, especially for most languages beyond English. Automated data collection offers a promising solution. This work presents a practical approach to generate AVSR datasets from raw video, refining existing techniques for improved efficiency and accessibility. We demonstrate its broad applicability by developing a baseline AVSR model for Vietnamese. Experiments show the automatically collected dataset enables a strong baseline, achieving competitive performance with robust ASR in clean conditions and significantly outperforming them in noisy environments like cocktail parties. This efficient method provides a pathway to expand AVSR to more languages, particularly under-resourced ones.
title ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.04635