Saved in:
Bibliographic Details
Main Authors: Lehečka, Jan, Psutka, Josef V., Šmídl, Luboš, Ircing, Pavel, Psutka, Josef
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2407.17160
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909324944605184
author Lehečka, Jan
Psutka, Josef V.
Šmídl, Luboš
Ircing, Pavel
Psutka, Josef
author_facet Lehečka, Jan
Psutka, Josef V.
Šmídl, Luboš
Ircing, Pavel
Psutka, Josef
contents In this paper, we are comparing monolingual Wav2Vec 2.0 models with various multilingual models to see whether we could improve speech recognition performance on a unique oral history archive containing a lot of mixed-language sentences. Our main goal is to push forward research on this unique dataset, which is an extremely valuable part of our cultural heritage. Our results suggest that monolingual speech recognition models are, in most cases, superior to multilingual models, even when processing the oral history archive full of mixed-language sentences from non-native speakers. We also performed the same experiments on the public CommonVoice dataset to verify our results. We are contributing to the research community by releasing our pre-trained models to the public.
format Preprint
id arxiv_https___arxiv_org_abs_2407_17160
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Comparative Analysis of Bilingual and Trilingual Wav2Vec Models for Automatic Speech Recognition in Multilingual Oral History Archives
Lehečka, Jan
Psutka, Josef V.
Šmídl, Luboš
Ircing, Pavel
Psutka, Josef
Computation and Language
Artificial Intelligence
In this paper, we are comparing monolingual Wav2Vec 2.0 models with various multilingual models to see whether we could improve speech recognition performance on a unique oral history archive containing a lot of mixed-language sentences. Our main goal is to push forward research on this unique dataset, which is an extremely valuable part of our cultural heritage. Our results suggest that monolingual speech recognition models are, in most cases, superior to multilingual models, even when processing the oral history archive full of mixed-language sentences from non-native speakers. We also performed the same experiments on the public CommonVoice dataset to verify our results. We are contributing to the research community by releasing our pre-trained models to the public.
title A Comparative Analysis of Bilingual and Trilingual Wav2Vec Models for Automatic Speech Recognition in Multilingual Oral History Archives
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2407.17160