Towards measuring fairness in speech recognition: Fair-Speech dataset

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Veliche, Irina-Elena, Huang, Zhuangqun, Kochaniyan, Vineeth Ayyat, Peng, Fuchun, Kalinli, Ozlem, Seltzer, Michael L.
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914921338372096
author Veliche, Irina-Elena
Huang, Zhuangqun
Kochaniyan, Vineeth Ayyat
Peng, Fuchun
Kalinli, Ozlem
Seltzer, Michael L.
author_facet Veliche, Irina-Elena
Huang, Zhuangqun
Kochaniyan, Vineeth Ayyat
Peng, Fuchun
Kalinli, Ozlem
Seltzer, Michael L.
contents The current public datasets for speech recognition (ASR) tend not to focus specifically on the fairness aspect, such as performance across different demographic groups. This paper introduces a novel dataset, Fair-Speech, a publicly released corpus to help researchers evaluate their ASR models for accuracy across a diverse set of self-reported demographic information, such as age, gender, ethnicity, geographic variation and whether the participants consider themselves native English speakers. Our dataset includes approximately 26.5K utterances in recorded speech by 593 people in the United States, who were paid to record and submit audios of themselves saying voice commands. We also provide ASR baselines, including on models trained on transcribed and untranscribed social media videos and open source models.
format Preprint
id arxiv_https___arxiv_org_abs_2408_12734
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Towards measuring fairness in speech recognition: Fair-Speech dataset
Veliche, Irina-Elena
Huang, Zhuangqun
Kochaniyan, Vineeth Ayyat
Peng, Fuchun
Kalinli, Ozlem
Seltzer, Michael L.
Artificial Intelligence
Computers and Society
Sound
Audio and Speech Processing
Machine Learning
The current public datasets for speech recognition (ASR) tend not to focus specifically on the fairness aspect, such as performance across different demographic groups. This paper introduces a novel dataset, Fair-Speech, a publicly released corpus to help researchers evaluate their ASR models for accuracy across a diverse set of self-reported demographic information, such as age, gender, ethnicity, geographic variation and whether the participants consider themselves native English speakers. Our dataset includes approximately 26.5K utterances in recorded speech by 593 people in the United States, who were paid to record and submit audios of themselves saying voice commands. We also provide ASR baselines, including on models trained on transcribed and untranscribed social media videos and open source models.
title Towards measuring fairness in speech recognition: Fair-Speech dataset
topic Artificial Intelligence
Computers and Society
Sound
Audio and Speech Processing
Machine Learning
url https://arxiv.org/abs/2408.12734