You don't understand me!: Comparing ASR results for L1 and L2 speakers of Swedish

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cumbal, Ronald, Moell, Birger, Lopes, Jose, Engwall, Olof
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909208953225216
author Cumbal, Ronald
Moell, Birger
Lopes, Jose
Engwall, Olof
author_facet Cumbal, Ronald
Moell, Birger
Lopes, Jose
Engwall, Olof
contents The performance of Automatic Speech Recognition (ASR) systems has constantly increased in state-of-the-art development. However, performance tends to decrease considerably in more challenging conditions (e.g., background noise, multiple speaker social conversations) and with more atypical speakers (e.g., children, non-native speakers or people with speech disorders), which signifies that general improvements do not necessarily transfer to applications that rely on ASR, e.g., educational software for younger students or language learners. In this study, we focus on the gap in performance between recognition results for native and non-native, read and spontaneous, Swedish utterances transcribed by different ASR services. We compare the recognition results using Word Error Rate and analyze the linguistic factors that may generate the observed transcription errors.
format Preprint
id arxiv_https___arxiv_org_abs_2405_13379
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle You don't understand me!: Comparing ASR results for L1 and L2 speakers of Swedish
Cumbal, Ronald
Moell, Birger
Lopes, Jose
Engwall, Olof
Computation and Language
Sound
Audio and Speech Processing
The performance of Automatic Speech Recognition (ASR) systems has constantly increased in state-of-the-art development. However, performance tends to decrease considerably in more challenging conditions (e.g., background noise, multiple speaker social conversations) and with more atypical speakers (e.g., children, non-native speakers or people with speech disorders), which signifies that general improvements do not necessarily transfer to applications that rely on ASR, e.g., educational software for younger students or language learners. In this study, we focus on the gap in performance between recognition results for native and non-native, read and spontaneous, Swedish utterances transcribed by different ASR services. We compare the recognition results using Word Error Rate and analyze the linguistic factors that may generate the observed transcription errors.
title You don't understand me!: Comparing ASR results for L1 and L2 speakers of Swedish
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2405.13379