ASR Benchmarking: Need for a More Representative Conversational Dataset
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913506423472128 |
|---|---|
| author | Maheshwari, Gaurav Ivanov, Dmitry Johannet, Théo Haddad, Kevin El |
| author_facet | Maheshwari, Gaurav Ivanov, Dmitry Johannet, Théo Haddad, Kevin El |
| contents | Automatic Speech Recognition (ASR) systems have achieved remarkable performance on widely used benchmarks such as LibriSpeech and Fleurs. However, these benchmarks do not adequately reflect the complexities of real-world conversational environments, where speech is often unstructured and contains disfluencies such as pauses, interruptions, and diverse accents. In this study, we introduce a multilingual conversational dataset, derived from TalkBank, consisting of unstructured phone conversation between adults. Our results show a significant performance drop across various state-of-the-art ASR models when tested in conversational settings. Furthermore, we observe a correlation between Word Error Rate and the presence of speech disfluencies, highlighting the critical need for more realistic, conversational ASR benchmarks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2409_12042 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | ASR Benchmarking: Need for a More Representative Conversational Dataset Maheshwari, Gaurav Ivanov, Dmitry Johannet, Théo Haddad, Kevin El Computation and Language Sound Audio and Speech Processing Automatic Speech Recognition (ASR) systems have achieved remarkable performance on widely used benchmarks such as LibriSpeech and Fleurs. However, these benchmarks do not adequately reflect the complexities of real-world conversational environments, where speech is often unstructured and contains disfluencies such as pauses, interruptions, and diverse accents. In this study, we introduce a multilingual conversational dataset, derived from TalkBank, consisting of unstructured phone conversation between adults. Our results show a significant performance drop across various state-of-the-art ASR models when tested in conversational settings. Furthermore, we observe a correlation between Word Error Rate and the presence of speech disfluencies, highlighting the critical need for more realistic, conversational ASR benchmarks. |
| title | ASR Benchmarking: Need for a More Representative Conversational Dataset |
| topic | Computation and Language Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2409.12042 |