Low-Resource, High-Impact: Building Corpora for Inclusive Language Technologies
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Artemova, Ekaterina, Burchell, Laurie, Dementieva, Daryna, Okabe, Shu, Shmatova, Mariya, Suarez, Pedro Ortiz |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MultiParaDetox: Extending Text Detoxification with Parallel Data to New Languages
von: Dementieva, Daryna, et al.
Veröffentlicht: (2024)
von: Dementieva, Daryna, et al.
Veröffentlicht: (2024)
Cross-lingual Text Classification Transfer: The Case of Ukrainian
von: Dementieva, Daryna, et al.
Veröffentlicht: (2024)
von: Dementieva, Daryna, et al.
Veröffentlicht: (2024)
Multilingual and Explainable Text Detoxification with Parallel Corpora
von: Dementieva, Daryna, et al.
Veröffentlicht: (2024)
von: Dementieva, Daryna, et al.
Veröffentlicht: (2024)
Temperature Matters: Enhancing Watermark Robustness Against Paraphrasing Attacks
von: Idrissi, Badr Youbi, et al.
Veröffentlicht: (2025)
von: Idrissi, Badr Youbi, et al.
Veröffentlicht: (2025)
GhanaNLP Parallel Corpora: Comprehensive Multilingual Resources for Low-Resource Ghanaian Languages
von: Gyamfi, Lawrence Adu, et al.
Veröffentlicht: (2026)
von: Gyamfi, Lawrence Adu, et al.
Veröffentlicht: (2026)
OpenWHO: A Document-Level Parallel Corpus for Health Translation in Low-Resource Languages
von: Merx, Raphaël, et al.
Veröffentlicht: (2025)
von: Merx, Raphaël, et al.
Veröffentlicht: (2025)
Can Prompting LLMs Unlock Hate Speech Detection across Languages? A Zero-shot and Few-shot Study
von: Ghorbanpour, Faeze, et al.
Veröffentlicht: (2025)
von: Ghorbanpour, Faeze, et al.
Veröffentlicht: (2025)
U-MATH: A University-Level Benchmark for Evaluating Mathematical Skills in LLMs
von: Chernyshev, Konstantin, et al.
Veröffentlicht: (2024)
von: Chernyshev, Konstantin, et al.
Veröffentlicht: (2024)
CrossNews-UA: A Cross-lingual News Semantic Similarity Benchmark for Ukrainian, Polish, Russian, and English
von: Dementieva, Daryna, et al.
Veröffentlicht: (2025)
von: Dementieva, Daryna, et al.
Veröffentlicht: (2025)
EmoBench-UA: A Benchmark Dataset for Emotion Detection in Ukrainian
von: Dementieva, Daryna, et al.
Veröffentlicht: (2025)
von: Dementieva, Daryna, et al.
Veröffentlicht: (2025)
JEEM: Vision-Language Understanding in Four Arabic Dialects
von: Kadaoui, Karima, et al.
Veröffentlicht: (2025)
von: Kadaoui, Karima, et al.
Veröffentlicht: (2025)
Evaluating Text Style Transfer: A Nine-Language Benchmark for Text Detoxification
von: Protasov, Vitaly, et al.
Veröffentlicht: (2025)
von: Protasov, Vitaly, et al.
Veröffentlicht: (2025)
Speech LLMs in Low-Resource Scenarios: Data Volume Requirements and the Impact of Pretraining on High-Resource Languages
von: Fong, Seraphina, et al.
Veröffentlicht: (2025)
von: Fong, Seraphina, et al.
Veröffentlicht: (2025)
Data-Efficient Hate Speech Detection via Cross-Lingual Nearest Neighbor Retrieval with Limited Labeled Data
von: Ghorbanpour, Faeze, et al.
Veröffentlicht: (2025)
von: Ghorbanpour, Faeze, et al.
Veröffentlicht: (2025)
Transcending Language Boundaries: Harnessing LLMs for Low-Resource Language Translation
von: Shu, Peng, et al.
Veröffentlicht: (2024)
von: Shu, Peng, et al.
Veröffentlicht: (2024)
Grounding Synthetic Data Evaluations of Language Models in Unsupervised Document Corpora
von: Majurski, Michael, et al.
Veröffentlicht: (2025)
von: Majurski, Michael, et al.
Veröffentlicht: (2025)
A Survey on Multilingual Large Language Models: Corpora, Alignment, and Bias
von: Xu, Yuemei, et al.
Veröffentlicht: (2024)
von: Xu, Yuemei, et al.
Veröffentlicht: (2024)
Multilingual Idioms in Sentences and Conversations Across High-, Medium-, and Low-Resource Languages
von: Almheiri, Saeed, et al.
Veröffentlicht: (2026)
von: Almheiri, Saeed, et al.
Veröffentlicht: (2026)
Mind the Language Gap: Automated and Augmented Evaluation of Bias in LLMs for High- and Low-Resource Languages
von: Buscemi, Alessio, et al.
Veröffentlicht: (2025)
von: Buscemi, Alessio, et al.
Veröffentlicht: (2025)
Explainable Semantic Textual Similarity via Dissimilar Span Detection
von: Lozano, Diego Miguel, et al.
Veröffentlicht: (2026)
von: Lozano, Diego Miguel, et al.
Veröffentlicht: (2026)
AI Brown and AI Koditex: LLM-Generated Corpora Comparable to Traditional Corpora of English and Czech Texts
von: Milička, Jiří, et al.
Veröffentlicht: (2025)
von: Milička, Jiří, et al.
Veröffentlicht: (2025)
Tokenization Matters: Navigating Data-Scarce Tokenization for Gender Inclusive Language Technologies
von: Ovalle, Anaelia, et al.
Veröffentlicht: (2023)
von: Ovalle, Anaelia, et al.
Veröffentlicht: (2023)
SaudiBERT: A Large Language Model Pretrained on Saudi Dialect Corpora
von: Qarah, Faisal
Veröffentlicht: (2024)
von: Qarah, Faisal
Veröffentlicht: (2024)
Guylingo: The Republic of Guyana Creole Corpora
von: Clarke, Christopher, et al.
Veröffentlicht: (2024)
von: Clarke, Christopher, et al.
Veröffentlicht: (2024)
Align and Shine: Building High-Quality Sentence-Aligned Corpora for Multilingual Text Simplification
von: Hilasaca, Kenji, et al.
Veröffentlicht: (2026)
von: Hilasaca, Kenji, et al.
Veröffentlicht: (2026)
Improving Semantic Understanding in Speech Language Models via Brain-tuning
von: Moussa, Omer, et al.
Veröffentlicht: (2024)
von: Moussa, Omer, et al.
Veröffentlicht: (2024)
Building a Chinese Medical Dialogue System: Integrating Large-scale Corpora and Novel Models
von: Wang, Xinyuan, et al.
Veröffentlicht: (2024)
von: Wang, Xinyuan, et al.
Veröffentlicht: (2024)
Dependency Parsing Across the Resource Spectrum: Evaluating Architectures on High and Low-Resource Languages
von: Guan, Kevin, et al.
Veröffentlicht: (2026)
von: Guan, Kevin, et al.
Veröffentlicht: (2026)
Transformers for Low-Resource Languages: Is Féidir Linn!
von: Lankford, Séamus, et al.
Veröffentlicht: (2024)
von: Lankford, Séamus, et al.
Veröffentlicht: (2024)
CorIL: Towards Enriching Indian Language to Indian Language Parallel Corpora and Machine Translation Systems
von: Bhattacharjee, Soham, et al.
Veröffentlicht: (2025)
von: Bhattacharjee, Soham, et al.
Veröffentlicht: (2025)
Attributing Culture-Conditioned Generations to Pretraining Corpora
von: Li, Huihan, et al.
Veröffentlicht: (2024)
von: Li, Huihan, et al.
Veröffentlicht: (2024)
Unlocking the Potential of Model Merging for Low-Resource Languages
von: Tao, Mingxu, et al.
Veröffentlicht: (2024)
von: Tao, Mingxu, et al.
Veröffentlicht: (2024)
BhashaSetu: Cross-Lingual Knowledge Transfer from High-Resource to Extreme Low-Resource Languages
von: Maji, Subhadip, et al.
Veröffentlicht: (2026)
von: Maji, Subhadip, et al.
Veröffentlicht: (2026)
Beyond Line-Level Filtering for the Pretraining Corpora of LLMs
von: Park, Chanwoo, et al.
Veröffentlicht: (2025)
von: Park, Chanwoo, et al.
Veröffentlicht: (2025)
Toxicity Classification in Ukrainian
von: Dementieva, Daryna, et al.
Veröffentlicht: (2024)
von: Dementieva, Daryna, et al.
Veröffentlicht: (2024)
Preference Consistency Matters: Enhancing Preference Learning in Language Models with Automated Self-Curation of Training Corpora
von: Lee, JoonHo, et al.
Veröffentlicht: (2024)
von: Lee, JoonHo, et al.
Veröffentlicht: (2024)
Continual-learning for Modelling Low-Resource Languages from Large Language Models
von: K, Santosh Srinath, et al.
Veröffentlicht: (2026)
von: K, Santosh Srinath, et al.
Veröffentlicht: (2026)
Opportunities and Challenges of Large Language Models for Low-Resource Languages in Humanities Research
von: Zhong, Tianyang, et al.
Veröffentlicht: (2024)
von: Zhong, Tianyang, et al.
Veröffentlicht: (2024)
LLMs Are Few-Shot In-Context Low-Resource Language Learners
von: Cahyawijaya, Samuel, et al.
Veröffentlicht: (2024)
von: Cahyawijaya, Samuel, et al.
Veröffentlicht: (2024)
Multilingual Transfer and Domain Adaptation for Low-Resource Languages of Spain
von: Luo, Yuanchang, et al.
Veröffentlicht: (2024)
von: Luo, Yuanchang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
MultiParaDetox: Extending Text Detoxification with Parallel Data to New Languages
von: Dementieva, Daryna, et al.
Veröffentlicht: (2024) -
Cross-lingual Text Classification Transfer: The Case of Ukrainian
von: Dementieva, Daryna, et al.
Veröffentlicht: (2024) -
Multilingual and Explainable Text Detoxification with Parallel Corpora
von: Dementieva, Daryna, et al.
Veröffentlicht: (2024) -
Temperature Matters: Enhancing Watermark Robustness Against Paraphrasing Attacks
von: Idrissi, Badr Youbi, et al.
Veröffentlicht: (2025) -
GhanaNLP Parallel Corpora: Comprehensive Multilingual Resources for Low-Resource Ghanaian Languages
von: Gyamfi, Lawrence Adu, et al.
Veröffentlicht: (2026)