Building low-resource African language corpora: A case study of Kidawida, Kalenjin and Dholuo
Fuente:
arXiv
Guardado en:
| Autores principales: | Mbogho, Audrey, Awuor, Quin, Kipkebut, Andrew, Wanzare, Lilian, Oloo, Vivian |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Kencorpus: A Kenyan Language Corpus of Swahili, Dholuo and Luhya for Natural Language Processing Tasks
por: Wanjawa, Barack, et al.
Publicado: (2022)
por: Wanjawa, Barack, et al.
Publicado: (2022)
AfriVoices-KE: A Multilingual Speech Dataset for Kenyan Languages
por: Wanzare, Lilian, et al.
Publicado: (2026)
por: Wanzare, Lilian, et al.
Publicado: (2026)
InkubaLM: A small language model for low-resource African languages
por: Tonja, Atnafu Lambebo, et al.
Publicado: (2024)
por: Tonja, Atnafu Lambebo, et al.
Publicado: (2024)
Yor-Sarc: A gold-standard dataset for sarcasm detection in a low-resource African language
por: Jimoh, Toheeb Aduramomi, et al.
Publicado: (2026)
por: Jimoh, Toheeb Aduramomi, et al.
Publicado: (2026)
Curating corpora with classifiers: A case study of clean energy sentiment online
por: Arnold, Michael V., et al.
Publicado: (2023)
por: Arnold, Michael V., et al.
Publicado: (2023)
Kenyan Sign Language (KSL) Dataset: Using Artificial Intelligence (AI) in Bridging Communication Barrier among the Deaf Learners
por: Wanzare, Lilian, et al.
Publicado: (2024)
por: Wanzare, Lilian, et al.
Publicado: (2024)
Effective vocabulary expanding of multilingual language models for extremely low-resource languages
por: Zheng, Jianyu
Publicado: (2026)
por: Zheng, Jianyu
Publicado: (2026)
Large corpora and large language models: a replicable method for automating grammatical annotation
por: Morin, Cameron, et al.
Publicado: (2024)
por: Morin, Cameron, et al.
Publicado: (2024)
Assessing effect sizes, variability, and power in the on-line study of language production
por: Audrey, Bürki, et al.
Publicado: (2024)
por: Audrey, Bürki, et al.
Publicado: (2024)
Multilingual corpora for the study of new concepts in the social sciences and humanities:
por: Kyriakoglou, Revekka, et al.
Publicado: (2025)
por: Kyriakoglou, Revekka, et al.
Publicado: (2025)
Corpora deduplication or duplication in Natural Language Processing of few resourced languages ? A case of study: The Mexico's Nahuatl
por: Guzman-Landa, Juan-José, et al.
Publicado: (2026)
por: Guzman-Landa, Juan-José, et al.
Publicado: (2026)
Cross-lingual neural fuzzy matching for exploiting target-language monolingual corpora in computer-aided translation
por: Esplà-Gomis, Miquel, et al.
Publicado: (2024)
por: Esplà-Gomis, Miquel, et al.
Publicado: (2024)
Phonetically rich corpus construction for a low-resourced language
por: Amadeus, Marcellus, et al.
Publicado: (2024)
por: Amadeus, Marcellus, et al.
Publicado: (2024)
Cross-lingual transfer of multilingual models on low resource African Languages
por: Thangaraj, Harish, et al.
Publicado: (2024)
por: Thangaraj, Harish, et al.
Publicado: (2024)
Multilingual jailbreaking of LLMs using low-resource languages
por: Marx, Dylan, et al.
Publicado: (2026)
por: Marx, Dylan, et al.
Publicado: (2026)
Prompt and circumstance: A word-by-word LLM prompting approach to interlinear glossing for low-resource languages
por: Elsner, Micha, et al.
Publicado: (2025)
por: Elsner, Micha, et al.
Publicado: (2025)
Are ASR foundation models generalized enough to capture features of regional dialects for low-resource languages?
por: Dipto, Tawsif Tashwar, et al.
Publicado: (2025)
por: Dipto, Tawsif Tashwar, et al.
Publicado: (2025)
Selective Attention Merging for low resource tasks: A case study of Child ASR
por: Shankar, Natarajan Balaji, et al.
Publicado: (2025)
por: Shankar, Natarajan Balaji, et al.
Publicado: (2025)
Two CFG Nahuatl for automatic corpora expansion
por: Guzmán-Landa, Juan-José, et al.
Publicado: (2025)
por: Guzmán-Landa, Juan-José, et al.
Publicado: (2025)
Leveraging LLMs for MT in Crisis Scenarios: a blueprint for low-resource languages
por: Lankford, Séamus, et al.
Publicado: (2024)
por: Lankford, Séamus, et al.
Publicado: (2024)
A comparison of pipelines for the translation of a low resource language based on transformers
por: Bonfanti, Chiara, et al.
Publicado: (2025)
por: Bonfanti, Chiara, et al.
Publicado: (2025)
Impact of enriched meaning representations for language generation in dialogue tasks: A comprehensive exploration of the relevance of tasks, corpora and metrics
por: Vázquez, Alain, et al.
Publicado: (2026)
por: Vázquez, Alain, et al.
Publicado: (2026)
Natural language processing for African languages
por: Adelani, David Ifeoluwa
Publicado: (2025)
por: Adelani, David Ifeoluwa
Publicado: (2025)
Development and bilingual evaluation of Japanese medical large language model within reasonably low computational resources
por: Sukeda, Issey
Publicado: (2024)
por: Sukeda, Issey
Publicado: (2024)
synthocr-gen: A synthetic ocr dataset generator for low-resource languages- breaking the data barrier
por: Malik, Haq Nawaz, et al.
Publicado: (2026)
por: Malik, Haq Nawaz, et al.
Publicado: (2026)
Under-resourced studies of under-resourced languages: lemmatization and POS-tagging with LLM annotators for historical Armenian, Georgian, Greek and Syriac
por: Vidal-Gorène, Chahan, et al.
Publicado: (2026)
por: Vidal-Gorène, Chahan, et al.
Publicado: (2026)
Data interference: emojis, homoglyphs, and issues of data fidelity in corpora and their results
por: Di Cristofaro, Matteo
Publicado: (2025)
por: Di Cristofaro, Matteo
Publicado: (2025)
Ukrainian-to-English folktale corpus: Parallel corpus creation and augmentation for machine translation in low-resource languages
por: Burda-Lassen, Olena
Publicado: (2024)
por: Burda-Lassen, Olena
Publicado: (2024)
KenSwQuAD -- A Question Answering Dataset for Swahili Low Resource Language
por: Wanjawa, Barack W., et al.
Publicado: (2022)
por: Wanjawa, Barack W., et al.
Publicado: (2022)
How do datasets, developers, and models affect biases in a low-resourced language?: The Case of the Bengali Language
por: Das, Dipto, et al.
Publicado: (2025)
por: Das, Dipto, et al.
Publicado: (2025)
Revisiting subword tokenization: A case study on affixal negation in large language models
por: Truong, Thinh Hung, et al.
Publicado: (2024)
por: Truong, Thinh Hung, et al.
Publicado: (2024)
The Greek podcast corpus: Competitive speech models for low-resourced languages with weakly supervised data
por: Paraskevopoulos, Georgios, et al.
Publicado: (2024)
por: Paraskevopoulos, Georgios, et al.
Publicado: (2024)
Entropy and type-token ratio in gigaword corpora
por: Rosillo-Rodes, Pablo, et al.
Publicado: (2024)
por: Rosillo-Rodes, Pablo, et al.
Publicado: (2024)
NER- RoBERTa: Fine-Tuning RoBERTa for Named Entity Recognition (NER) within low-resource languages
por: Abdullah, Abdulhady Abas, et al.
Publicado: (2024)
por: Abdullah, Abdulhady Abas, et al.
Publicado: (2024)
DHPLT: large-scale multilingual diachronic corpora and word representations for semantic change modelling
por: Fedorova, Mariia, et al.
Publicado: (2026)
por: Fedorova, Mariia, et al.
Publicado: (2026)
How language models extrapolate outside the training data: A case study in Textualized Gridworld
por: Kim, Doyoung, et al.
Publicado: (2024)
por: Kim, Doyoung, et al.
Publicado: (2024)
UstanceBR: a social media language resource for stance prediction
por: Pereira, Camila, et al.
Publicado: (2023)
por: Pereira, Camila, et al.
Publicado: (2023)
Dynamic Embedded Topic Models: properties and recommendations based on diverse corpora
por: Fittschen, Elisabeth, et al.
Publicado: (2025)
por: Fittschen, Elisabeth, et al.
Publicado: (2025)
Language corpora for the Dutch medical domain
por: van Es, B.
Publicado: (2026)
por: van Es, B.
Publicado: (2026)
Strategic resource allocation in memory encoding: An efficiency principle shaping language processing
por: Xu, Weijie, et al.
Publicado: (2025)
por: Xu, Weijie, et al.
Publicado: (2025)
Ejemplares similares
-
Kencorpus: A Kenyan Language Corpus of Swahili, Dholuo and Luhya for Natural Language Processing Tasks
por: Wanjawa, Barack, et al.
Publicado: (2022) -
AfriVoices-KE: A Multilingual Speech Dataset for Kenyan Languages
por: Wanzare, Lilian, et al.
Publicado: (2026) -
InkubaLM: A small language model for low-resource African languages
por: Tonja, Atnafu Lambebo, et al.
Publicado: (2024) -
Yor-Sarc: A gold-standard dataset for sarcasm detection in a low-resource African language
por: Jimoh, Toheeb Aduramomi, et al.
Publicado: (2026) -
Curating corpora with classifiers: A case study of clean energy sentiment online
por: Arnold, Michael V., et al.
Publicado: (2023)