Efficiently Identifying Low-Quality Language Subsets in Multilingual Datasets: A Case Study on a Large-Scale Multilingual Audio Dataset

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Samir, Farhan, Ahn, Emily P., Prakash, Shreya, Soskuthy, Márton, Shwartz, Vered, Zhu, Jian
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916424648228864
author Samir, Farhan
Ahn, Emily P.
Prakash, Shreya
Soskuthy, Márton
Shwartz, Vered
Zhu, Jian
author_facet Samir, Farhan
Ahn, Emily P.
Prakash, Shreya
Soskuthy, Márton
Shwartz, Vered
Zhu, Jian
contents Curating datasets that span multiple languages is challenging. To make the collection more scalable, researchers often incorporate one or more imperfect classifiers in the process, like language identification models. These models, however, are prone to failure, resulting in some language subsets being unreliable for downstream tasks. We introduce a statistical test, the Preference Proportion Test, for identifying such unreliable subsets. By annotating only 20 samples for a language subset, we're able to identify systematic transcription errors for 10 language subsets in a recent large multilingual transcribed audio dataset, X-IPAPack (Zhu et al., 2024). We find that filtering this low-quality data out when training models for the downstream task of phonetic transcription brings substantial benefits, most notably a 25.7% relative improvement on transcribing recordings in out-of-distribution languages. Our method lays a path forward for systematic and reliable multilingual dataset auditing.
format Preprint
id arxiv_https___arxiv_org_abs_2410_04292
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Efficiently Identifying Low-Quality Language Subsets in Multilingual Datasets: A Case Study on a Large-Scale Multilingual Audio Dataset
Samir, Farhan
Ahn, Emily P.
Prakash, Shreya
Soskuthy, Márton
Shwartz, Vered
Zhu, Jian
Computation and Language
Curating datasets that span multiple languages is challenging. To make the collection more scalable, researchers often incorporate one or more imperfect classifiers in the process, like language identification models. These models, however, are prone to failure, resulting in some language subsets being unreliable for downstream tasks. We introduce a statistical test, the Preference Proportion Test, for identifying such unreliable subsets. By annotating only 20 samples for a language subset, we're able to identify systematic transcription errors for 10 language subsets in a recent large multilingual transcribed audio dataset, X-IPAPack (Zhu et al., 2024). We find that filtering this low-quality data out when training models for the downstream task of phonetic transcription brings substantial benefits, most notably a 25.7% relative improvement on transcribing recordings in out-of-distribution languages. Our method lays a path forward for systematic and reliable multilingual dataset auditing.
title Efficiently Identifying Low-Quality Language Subsets in Multilingual Datasets: A Case Study on a Large-Scale Multilingual Audio Dataset
topic Computation and Language
url https://arxiv.org/abs/2410.04292