On the Role of Speech Data in Reducing Toxicity Detection Bias

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Bell, Samuel J., Meglioli, Mariano Coria, Richards, Megan, Sánchez, Eduardo, Ropers, Christophe, Wang, Skyler, Williams, Adina, Sagun, Levent, Costa-jussà, Marta R.
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913840332013568
author Bell, Samuel J.
Meglioli, Mariano Coria
Richards, Megan
Sánchez, Eduardo
Ropers, Christophe
Wang, Skyler
Williams, Adina
Sagun, Levent
Costa-jussà, Marta R.
author_facet Bell, Samuel J.
Meglioli, Mariano Coria
Richards, Megan
Sánchez, Eduardo
Ropers, Christophe
Wang, Skyler
Williams, Adina
Sagun, Levent
Costa-jussà, Marta R.
contents Text toxicity detection systems exhibit significant biases, producing disproportionate rates of false positives on samples mentioning demographic groups. But what about toxicity detection in speech? To investigate the extent to which text-based biases are mitigated by speech-based systems, we produce a set of high-quality group annotations for the multilingual MuTox dataset, and then leverage these annotations to systematically compare speech- and text-based toxicity classifiers. Our findings indicate that access to speech data during inference supports reduced bias against group mentions, particularly for ambiguous and disagreement-inducing samples. Our results also suggest that improving classifiers, rather than transcription pipelines, is more helpful for reducing group bias. We publicly release our annotations and provide recommendations for future toxicity dataset construction.
format Preprint
id arxiv_https___arxiv_org_abs_2411_08135
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle On the Role of Speech Data in Reducing Toxicity Detection Bias
Bell, Samuel J.
Meglioli, Mariano Coria
Richards, Megan
Sánchez, Eduardo
Ropers, Christophe
Wang, Skyler
Williams, Adina
Sagun, Levent
Costa-jussà, Marta R.
Computation and Language
Artificial Intelligence
Machine Learning
Sound
Audio and Speech Processing
Text toxicity detection systems exhibit significant biases, producing disproportionate rates of false positives on samples mentioning demographic groups. But what about toxicity detection in speech? To investigate the extent to which text-based biases are mitigated by speech-based systems, we produce a set of high-quality group annotations for the multilingual MuTox dataset, and then leverage these annotations to systematically compare speech- and text-based toxicity classifiers. Our findings indicate that access to speech data during inference supports reduced bias against group mentions, particularly for ambiguous and disagreement-inducing samples. Our results also suggest that improving classifiers, rather than transcription pipelines, is more helpful for reducing group bias. We publicly release our annotations and provide recommendations for future toxicity dataset construction.
title On the Role of Speech Data in Reducing Toxicity Detection Bias
topic Computation and Language
Artificial Intelligence
Machine Learning
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2411.08135