Late fusion ensembles for speech recognition on diverse input audio representations

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Jezidžić, Marin, Mihelčić, Matej
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912141588561920
author Jezidžić, Marin
Mihelčić, Matej
author_facet Jezidžić, Marin
Mihelčić, Matej
contents We explore diverse representations of speech audio, and their effect on a performance of late fusion ensemble of E-Branchformer models, applied to Automatic Speech Recognition (ASR) task. Although it is generally known that ensemble methods often improve the performance of the system even for speech recognition, it is very interesting to explore how ensembles of complex state-of-the-art models, such as medium-sized and large E-Branchformers, cope in this setting when their base models are trained on diverse representations of the input speech audio. The results are evaluated on four widely-used benchmark datasets: \textit{Librispeech, Aishell, Gigaspeech}, \textit{TEDLIUMv2} and show that improvements of $1\% - 14\%$ can still be achieved over the state-of-the-art models trained using comparable techniques on these datasets. A noteworthy observation is that such ensemble offers improvements even with the use of language models, although the gap is closing.
format Preprint
id arxiv_https___arxiv_org_abs_2412_01861
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Late fusion ensembles for speech recognition on diverse input audio representations
Jezidžić, Marin
Mihelčić, Matej
Audio and Speech Processing
Machine Learning
Sound
We explore diverse representations of speech audio, and their effect on a performance of late fusion ensemble of E-Branchformer models, applied to Automatic Speech Recognition (ASR) task. Although it is generally known that ensemble methods often improve the performance of the system even for speech recognition, it is very interesting to explore how ensembles of complex state-of-the-art models, such as medium-sized and large E-Branchformers, cope in this setting when their base models are trained on diverse representations of the input speech audio. The results are evaluated on four widely-used benchmark datasets: \textit{Librispeech, Aishell, Gigaspeech}, \textit{TEDLIUMv2} and show that improvements of $1\% - 14\%$ can still be achieved over the state-of-the-art models trained using comparable techniques on these datasets. A noteworthy observation is that such ensemble offers improvements even with the use of language models, although the gap is closing.
title Late fusion ensembles for speech recognition on diverse input audio representations
topic Audio and Speech Processing
Machine Learning
Sound
url https://arxiv.org/abs/2412.01861