How Good is Your Wikipedia? Auditing Data Quality for Low-resource and Multilingual NLP
Fuente:
arXiv
Saved in:
| Main Authors: | Tatariya, Kushal, Kulmizev, Artur, Poelman, Wessel, Ploeger, Esther, Bollmann, Marcel, Bjerva, Johannes, Luo, Jiaming, Lent, Heather, de Lhoneux, Miryam |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
What is "Typological Diversity" in NLP?
by: Ploeger, Esther, et al.
Published: (2024)
by: Ploeger, Esther, et al.
Published: (2024)
Sociolinguistically Informed Interpretability: A Case Study on Hinglish Emotion Classification
by: Tatariya, Kushal, et al.
Published: (2024)
by: Tatariya, Kushal, et al.
Published: (2024)
On the Interplay between Positional Encodings, Morphological Complexity, and Word Order Flexibility
by: Tatariya, Kushal, et al.
Published: (2025)
by: Tatariya, Kushal, et al.
Published: (2025)
Form and Meaning in Intrinsic Multilingual Evaluations
by: Poelman, Wessel, et al.
Published: (2026)
by: Poelman, Wessel, et al.
Published: (2026)
The Roles of English in Evaluating Multilingual Language Models
by: Poelman, Wessel, et al.
Published: (2024)
by: Poelman, Wessel, et al.
Published: (2024)
A Principled Framework for Evaluating on Typologically Diverse Languages
by: Ploeger, Esther, et al.
Published: (2024)
by: Ploeger, Esther, et al.
Published: (2024)
QQ: A Toolkit for Language Identifiers and Metadata
by: Poelman, Wessel, et al.
Published: (2026)
by: Poelman, Wessel, et al.
Published: (2026)
Confounding Factors in Relating Model Performance to Morphology
by: Poelman, Wessel, et al.
Published: (2025)
by: Poelman, Wessel, et al.
Published: (2025)
Pixology: Probing the Linguistic and Visual Capabilities of Pixel-based Language Models
by: Tatariya, Kushal, et al.
Published: (2024)
by: Tatariya, Kushal, et al.
Published: (2024)
Multilingual Gradient Word-Order Typology from Universal Dependencies
by: Baylor, Emi, et al.
Published: (2024)
by: Baylor, Emi, et al.
Published: (2024)
Text Embedding Inversion Security for Multilingual Language Models
by: Chen, Yiyi, et al.
Published: (2024)
by: Chen, Yiyi, et al.
Published: (2024)
CreoleVal: Multilingual Multitask Benchmarks for Creoles
by: Lent, Heather, et al.
Published: (2023)
by: Lent, Heather, et al.
Published: (2023)
Against All Odds: Overcoming Typology, Script, and Language Confusion in Multilingual Embedding Inversion Attacks
by: Chen, Yiyi, et al.
Published: (2024)
by: Chen, Yiyi, et al.
Published: (2024)
NLP Security and Ethics, in the Wild
by: Lent, Heather, et al.
Published: (2025)
by: Lent, Heather, et al.
Published: (2025)
Beyond Weaponization: NLP Security for Medium and Lower-Resourced Languages in Their Own Right
by: Lent, Heather
Published: (2025)
by: Lent, Heather
Published: (2025)
Type and Complexity Signals in Multilingual Question Representations
by: Kokot, Robin, et al.
Published: (2025)
by: Kokot, Robin, et al.
Published: (2025)
We Need to Measure Data Diversity in NLP -- Better and Broader
by: Nguyen, Dong, et al.
Published: (2025)
by: Nguyen, Dong, et al.
Published: (2025)
Connecting Ideas in 'Lower-Resource' Scenarios: NLP for National Varieties, Creoles and Other Low-resource Scenarios
by: Joshi, Aditya, et al.
Published: (2024)
by: Joshi, Aditya, et al.
Published: (2024)
Typologically Informed Parameter Aggregation
by: Accou, Stef, et al.
Published: (2026)
by: Accou, Stef, et al.
Published: (2026)
Trans-Tokenization and Cross-lingual Vocabulary Transfers: Language Adaptation of LLMs for Low-Resource NLP
by: Remy, François, et al.
Published: (2024)
by: Remy, François, et al.
Published: (2024)
Knowledge Graphs, Large Language Models, and Hallucinations: An NLP Perspective
by: Lavrinovics, Ernests, et al.
Published: (2024)
by: Lavrinovics, Ernests, et al.
Published: (2024)
Recipe for Zero-shot POS Tagging: Is It Useful in Realistic Scenarios?
by: Vandenbulcke, Zeno, et al.
Published: (2024)
by: Vandenbulcke, Zeno, et al.
Published: (2024)
Limited-Resource Adapters Are Regularizers, Not Linguists
by: Fekete, Marcell, et al.
Published: (2025)
by: Fekete, Marcell, et al.
Published: (2025)
Shared Path: Unraveling Memorization in Multilingual LLMs through Language Similarities
by: Luo, Xiaoyu, et al.
Published: (2025)
by: Luo, Xiaoyu, et al.
Published: (2025)
MultiHal: Multilingual Dataset for Knowledge-Graph Grounded Evaluation of LLM Hallucinations
by: Lavrinovics, Ernests, et al.
Published: (2025)
by: Lavrinovics, Ernests, et al.
Published: (2025)
Engineering Conversational Search Systems: A Review of Applications, Architectures, and Functional Components
by: Schneider, Phillip, et al.
Published: (2024)
by: Schneider, Phillip, et al.
Published: (2024)
Patterns of Persistence and Diffusibility across the World's Languages
by: Chen, Yiyi, et al.
Published: (2024)
by: Chen, Yiyi, et al.
Published: (2024)
Linguistically Grounded Analysis of Language Models using Shapley Head Values
by: Fekete, Marcell, et al.
Published: (2024)
by: Fekete, Marcell, et al.
Published: (2024)
Wikipedia Citations: Reproducible Citation Extraction from Multilingual Wikipedia
by: Kokash, Natallia, et al.
Published: (2024)
by: Kokash, Natallia, et al.
Published: (2024)
Signal Quality Auditing for Time-series Data
by: Gao, Chufan, et al.
Published: (2024)
by: Gao, Chufan, et al.
Published: (2024)
Factual Inconsistencies in Multilingual Wikipedia Tables
by: Cappa, Silvia, et al.
Published: (2025)
by: Cappa, Silvia, et al.
Published: (2025)
Libraries and Librarians as Depicted in Freshmen English Textbooks: An Update.
by: Lent, John
Published: (1991)
by: Lent, John
Published: (1991)
Unexpected Knowledge: Auditing Wikipedia and Grokipedia Search Recommendations
by: Coppolillo, Erica, et al.
Published: (2025)
by: Coppolillo, Erica, et al.
Published: (2025)
ALGEN: Few-shot Inversion Attacks on Textual Embeddings using Alignment and Generation
by: Chen, Yiyi, et al.
Published: (2025)
by: Chen, Yiyi, et al.
Published: (2025)
Follow the Path: Reasoning over Knowledge Graph Paths to Improve Large Language Model Factuality
by: Zhang, Mike, et al.
Published: (2025)
by: Zhang, Mike, et al.
Published: (2025)
When Discourse Pressures Conflict: Information Structure in Vision-Language Model Outputs
by: Fekete, Marcell, et al.
Published: (2026)
by: Fekete, Marcell, et al.
Published: (2026)
Multilingual Reference Need Assessment System for Wikipedia
by: Baigutanova, Aitolkyn, et al.
Published: (2026)
by: Baigutanova, Aitolkyn, et al.
Published: (2026)
An Open Multilingual System for Scoring Readability of Wikipedia
by: Trokhymovych, Mykola, et al.
Published: (2024)
by: Trokhymovych, Mykola, et al.
Published: (2024)
An Audit on the Perspectives and Challenges of Hallucinations in NLP
by: Venkit, Pranav Narayanan, et al.
Published: (2024)
by: Venkit, Pranav Narayanan, et al.
Published: (2024)
Controlling the Cascade: Kinematic Planning for N-ball Toss Juggling
by: Ploeger, Kai, et al.
Published: (2022)
by: Ploeger, Kai, et al.
Published: (2022)
Similar Items
-
What is "Typological Diversity" in NLP?
by: Ploeger, Esther, et al.
Published: (2024) -
Sociolinguistically Informed Interpretability: A Case Study on Hinglish Emotion Classification
by: Tatariya, Kushal, et al.
Published: (2024) -
On the Interplay between Positional Encodings, Morphological Complexity, and Word Order Flexibility
by: Tatariya, Kushal, et al.
Published: (2025) -
Form and Meaning in Intrinsic Multilingual Evaluations
by: Poelman, Wessel, et al.
Published: (2026) -
The Roles of English in Evaluating Multilingual Language Models
by: Poelman, Wessel, et al.
Published: (2024)