Low-Resource, High-Impact: Building Corpora for Inclusive Language Technologies

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Artemova, Ekaterina, Burchell, Laurie, Dementieva, Daryna, Okabe, Shu, Shmatova, Mariya, Suarez, Pedro Ortiz
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911322953744384
author Artemova, Ekaterina
Burchell, Laurie
Dementieva, Daryna
Okabe, Shu
Shmatova, Mariya
Suarez, Pedro Ortiz
author_facet Artemova, Ekaterina
Burchell, Laurie
Dementieva, Daryna
Okabe, Shu
Shmatova, Mariya
Suarez, Pedro Ortiz
contents This tutorial (https://tum-nlp.github.io/low-resource-tutorial) is designed for NLP practitioners, researchers, and developers working with multilingual and low-resource languages who seek to create more equitable and socially impactful language technologies. Participants will walk away with a practical toolkit for building end-to-end NLP pipelines for underrepresented languages -- from data collection and web crawling to parallel sentence mining, machine translation, and downstream applications such as text classification and multimodal reasoning. The tutorial presents strategies for tackling the challenges of data scarcity and cultural variance, offering hands-on methods and modeling frameworks. We will focus on fair, reproducible, and community-informed development approaches, grounded in real-world scenarios. We will showcase a diverse set of use cases covering over 10 languages from different language families and geopolitical contexts, including both digitally resource-rich and severely underrepresented languages.
format Preprint
id arxiv_https___arxiv_org_abs_2512_14576
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Low-Resource, High-Impact: Building Corpora for Inclusive Language Technologies
Artemova, Ekaterina
Burchell, Laurie
Dementieva, Daryna
Okabe, Shu
Shmatova, Mariya
Suarez, Pedro Ortiz
Computation and Language
Artificial Intelligence
This tutorial (https://tum-nlp.github.io/low-resource-tutorial) is designed for NLP practitioners, researchers, and developers working with multilingual and low-resource languages who seek to create more equitable and socially impactful language technologies. Participants will walk away with a practical toolkit for building end-to-end NLP pipelines for underrepresented languages -- from data collection and web crawling to parallel sentence mining, machine translation, and downstream applications such as text classification and multimodal reasoning. The tutorial presents strategies for tackling the challenges of data scarcity and cultural variance, offering hands-on methods and modeling frameworks. We will focus on fair, reproducible, and community-informed development approaches, grounded in real-world scenarios. We will showcase a diverse set of use cases covering over 10 languages from different language families and geopolitical contexts, including both digitally resource-rich and severely underrepresented languages.
title Low-Resource, High-Impact: Building Corpora for Inclusive Language Technologies
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2512.14576