Towards measuring fairness in speech recognition: Fair-Speech dataset
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866914921338372096 |
|---|---|
| author | Veliche, Irina-Elena Huang, Zhuangqun Kochaniyan, Vineeth Ayyat Peng, Fuchun Kalinli, Ozlem Seltzer, Michael L. |
| author_facet | Veliche, Irina-Elena Huang, Zhuangqun Kochaniyan, Vineeth Ayyat Peng, Fuchun Kalinli, Ozlem Seltzer, Michael L. |
| contents | The current public datasets for speech recognition (ASR) tend not to focus specifically on the fairness aspect, such as performance across different demographic groups. This paper introduces a novel dataset, Fair-Speech, a publicly released corpus to help researchers evaluate their ASR models for accuracy across a diverse set of self-reported demographic information, such as age, gender, ethnicity, geographic variation and whether the participants consider themselves native English speakers. Our dataset includes approximately 26.5K utterances in recorded speech by 593 people in the United States, who were paid to record and submit audios of themselves saying voice commands. We also provide ASR baselines, including on models trained on transcribed and untranscribed social media videos and open source models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2408_12734 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Towards measuring fairness in speech recognition: Fair-Speech dataset Veliche, Irina-Elena Huang, Zhuangqun Kochaniyan, Vineeth Ayyat Peng, Fuchun Kalinli, Ozlem Seltzer, Michael L. Artificial Intelligence Computers and Society Sound Audio and Speech Processing Machine Learning The current public datasets for speech recognition (ASR) tend not to focus specifically on the fairness aspect, such as performance across different demographic groups. This paper introduces a novel dataset, Fair-Speech, a publicly released corpus to help researchers evaluate their ASR models for accuracy across a diverse set of self-reported demographic information, such as age, gender, ethnicity, geographic variation and whether the participants consider themselves native English speakers. Our dataset includes approximately 26.5K utterances in recorded speech by 593 people in the United States, who were paid to record and submit audios of themselves saying voice commands. We also provide ASR baselines, including on models trained on transcribed and untranscribed social media videos and open source models. |
| title | Towards measuring fairness in speech recognition: Fair-Speech dataset |
| topic | Artificial Intelligence Computers and Society Sound Audio and Speech Processing Machine Learning |
| url | https://arxiv.org/abs/2408.12734 |