BERSting at the Screams: A Benchmark for Distanced, Emotional and Shouted Speech Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tuttösí, Paige, Dhillon, Mantaj, Sang, Luna, Eastwood, Shane, Bhatia, Poorvi, Dinh, Quang Minh, Kapoor, Avni, Jin, Yewon, Lim, Angelica
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912510156734464
author Tuttösí, Paige
Dhillon, Mantaj
Sang, Luna
Eastwood, Shane
Bhatia, Poorvi
Dinh, Quang Minh
Kapoor, Avni
Jin, Yewon
Lim, Angelica
author_facet Tuttösí, Paige
Dhillon, Mantaj
Sang, Luna
Eastwood, Shane
Bhatia, Poorvi
Dinh, Quang Minh
Kapoor, Avni
Jin, Yewon
Lim, Angelica
contents Some speech recognition tasks, such as automatic speech recognition (ASR), are approaching or have reached human performance in many reported metrics. Yet, they continue to struggle in complex, real-world, situations, such as with distanced speech. Previous challenges have released datasets to address the issue of distanced ASR, however, the focus remains primarily on distance, specifically relying on multi-microphone array systems. Here we present the B(asic) E(motion) R(andom phrase) S(hou)t(s) (BERSt) dataset. The dataset contains almost 4 hours of English speech from 98 actors with varying regional and non-native accents. The data was collected on smartphones in the actors homes and therefore includes at least 98 different acoustic environments. The data also includes 7 different emotion prompts and both shouted and spoken utterances. The smartphones were places in 19 different positions, including obstructions and being in a different room than the actor. This data is publicly available for use and can be used to evaluate a variety of speech recognition tasks, including: ASR, shout detection, and speech emotion recognition (SER). We provide initial benchmarks for ASR and SER tasks, and find that ASR degrades both with an increase in distance and shout level and shows varied performance depending on the intended emotion. Our results show that the BERSt dataset is challenging for both ASR and SER tasks and continued work is needed to improve the robustness of such systems for more accurate real-world use.
format Preprint
id arxiv_https___arxiv_org_abs_2505_00059
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle BERSting at the Screams: A Benchmark for Distanced, Emotional and Shouted Speech Recognition
Tuttösí, Paige
Dhillon, Mantaj
Sang, Luna
Eastwood, Shane
Bhatia, Poorvi
Dinh, Quang Minh
Kapoor, Avni
Jin, Yewon
Lim, Angelica
Computation and Language
Sound
Audio and Speech Processing
Some speech recognition tasks, such as automatic speech recognition (ASR), are approaching or have reached human performance in many reported metrics. Yet, they continue to struggle in complex, real-world, situations, such as with distanced speech. Previous challenges have released datasets to address the issue of distanced ASR, however, the focus remains primarily on distance, specifically relying on multi-microphone array systems. Here we present the B(asic) E(motion) R(andom phrase) S(hou)t(s) (BERSt) dataset. The dataset contains almost 4 hours of English speech from 98 actors with varying regional and non-native accents. The data was collected on smartphones in the actors homes and therefore includes at least 98 different acoustic environments. The data also includes 7 different emotion prompts and both shouted and spoken utterances. The smartphones were places in 19 different positions, including obstructions and being in a different room than the actor. This data is publicly available for use and can be used to evaluate a variety of speech recognition tasks, including: ASR, shout detection, and speech emotion recognition (SER). We provide initial benchmarks for ASR and SER tasks, and find that ASR degrades both with an increase in distance and shout level and shows varied performance depending on the intended emotion. Our results show that the BERSt dataset is challenging for both ASR and SER tasks and continued work is needed to improve the robustness of such systems for more accurate real-world use.
title BERSting at the Screams: A Benchmark for Distanced, Emotional and Shouted Speech Recognition
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2505.00059