How Good is Your Wikipedia? Auditing Data Quality for Low-resource and Multilingual NLP

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tatariya, Kushal, Kulmizev, Artur, Poelman, Wessel, Ploeger, Esther, Bollmann, Marcel, Bjerva, Johannes, Luo, Jiaming, Lent, Heather, de Lhoneux, Miryam
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910185956573184
author Tatariya, Kushal
Kulmizev, Artur
Poelman, Wessel
Ploeger, Esther
Bollmann, Marcel
Bjerva, Johannes
Luo, Jiaming
Lent, Heather
de Lhoneux, Miryam
author_facet Tatariya, Kushal
Kulmizev, Artur
Poelman, Wessel
Ploeger, Esther
Bollmann, Marcel
Bjerva, Johannes
Luo, Jiaming
Lent, Heather
de Lhoneux, Miryam
contents Wikipedia's perceived high quality and broad language coverage have established it as a fundamental resource in NLP. However, in recent years, such assumptions of high quality have become the subject of scrutiny in low-resource and multilingual contexts. In this study, we subject the entirety of non-English Wikipedia to a data filtering procedure typically reserved for noisy web-text -- a process which removes a large percentage of the collection's data. In analysing the removed data, we reveal numerous systematic quality issues, such as script and language contamination, repeated template and placeholder articles, and a high concentration of bot-generated content. We consolidate these findings into a 4-level quality ranking of Wikipedia, which shows strong correspondence with alternative quality measures and heuristics. Lastly, we evaluate the downstream impact of quality filtering in three practical language modelling scenarios, showing that models trained on filtered data largely match or outperform those trained on raw Wikipedia, with the largest gains observed for lower-quality language editions. Ultimately, our experiments serve as a first step in establishing quality-aware best practices for Wikipedia utilization in NLP, laying groundwork that can inform future dataset creation and curation efforts.
format Preprint
id arxiv_https___arxiv_org_abs_2411_05527
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle How Good is Your Wikipedia? Auditing Data Quality for Low-resource and Multilingual NLP
Tatariya, Kushal
Kulmizev, Artur
Poelman, Wessel
Ploeger, Esther
Bollmann, Marcel
Bjerva, Johannes
Luo, Jiaming
Lent, Heather
de Lhoneux, Miryam
Computation and Language
Wikipedia's perceived high quality and broad language coverage have established it as a fundamental resource in NLP. However, in recent years, such assumptions of high quality have become the subject of scrutiny in low-resource and multilingual contexts. In this study, we subject the entirety of non-English Wikipedia to a data filtering procedure typically reserved for noisy web-text -- a process which removes a large percentage of the collection's data. In analysing the removed data, we reveal numerous systematic quality issues, such as script and language contamination, repeated template and placeholder articles, and a high concentration of bot-generated content. We consolidate these findings into a 4-level quality ranking of Wikipedia, which shows strong correspondence with alternative quality measures and heuristics. Lastly, we evaluate the downstream impact of quality filtering in three practical language modelling scenarios, showing that models trained on filtered data largely match or outperform those trained on raw Wikipedia, with the largest gains observed for lower-quality language editions. Ultimately, our experiments serve as a first step in establishing quality-aware best practices for Wikipedia utilization in NLP, laying groundwork that can inform future dataset creation and curation efforts.
title How Good is Your Wikipedia? Auditing Data Quality for Low-resource and Multilingual NLP
topic Computation and Language
url https://arxiv.org/abs/2411.05527