The Thiomi Dataset: A Large-Scale Multimodal Corpus for Low-Resource African Languages
Fuente:
arXiv
Guardado en:
| Autores principales: | Mutisya, Hillary, Mugane, John, Nyamboga, Gavin, Chege, Brian, Gathoni, Maryruth |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Zero-Shot Morphological Discovery in Low-Resource Bantu Languages via Cross-Lingual Transfer and Unsupervised Clustering
por: Mutisya, Hillary, et al.
Publicado: (2026)
por: Mutisya, Hillary, et al.
Publicado: (2026)
Neural Recovery of Historical Lexical Structure in Bantu Languages from Modern Data
por: Mutisya, Hillary, et al.
Publicado: (2026)
por: Mutisya, Hillary, et al.
Publicado: (2026)
Attention Sinks in Massively Multilingual Neural Machine Translation:Discovery, Analysis, and Mitigation
por: Mutisya, Hillary, et al.
Publicado: (2026)
por: Mutisya, Hillary, et al.
Publicado: (2026)
Continued Pretraining for Low-Resource Swahili ASR: Achieving State-of-the-Art Performance with Minimal Labeled Data
por: Mutisya, Hillary, et al.
Publicado: (2026)
por: Mutisya, Hillary, et al.
Publicado: (2026)
A Hybrid Protocol for Large-Scale Semantic Dataset Generation in Low-Resource Languages: The Turkish Semantic Relations Corpus
por: Tosun, Ebubekir, et al.
Publicado: (2026)
por: Tosun, Ebubekir, et al.
Publicado: (2026)
Large Multimodal Models for Low-Resource Languages: A Survey
por: Lupascu, Marian, et al.
Publicado: (2025)
por: Lupascu, Marian, et al.
Publicado: (2025)
Adapting Multilingual LLMs to Low-Resource Languages using Continued Pre-training and Synthetic Corpus
por: Joshi, Raviraj, et al.
Publicado: (2024)
por: Joshi, Raviraj, et al.
Publicado: (2024)
Large Language Models for Math Education in Low-Resource Languages: A Study in Sinhala and Tamil
por: Kishanthan, Sukumar, et al.
Publicado: (2026)
por: Kishanthan, Sukumar, et al.
Publicado: (2026)
SynDARin: Synthesising Datasets for Automated Reasoning in Low-Resource Languages
por: Ghazaryan, Gayane, et al.
Publicado: (2024)
por: Ghazaryan, Gayane, et al.
Publicado: (2024)
OceanPile: A Large-Scale Multimodal Ocean Corpus for Foundation Models
por: Xue, Yida, et al.
Publicado: (2026)
por: Xue, Yida, et al.
Publicado: (2026)
DACTYL: Diverse Adversarial Corpus of Texts Yielded from Large Language Models
por: Thorat, Shantanu, et al.
Publicado: (2025)
por: Thorat, Shantanu, et al.
Publicado: (2025)
PashtoCorp: A 1.25-Billion-Word Corpus, Evaluation Suite, and Reproducible Pipeline for Low-Resource Language Development
por: Rahman, Hanif
Publicado: (2026)
por: Rahman, Hanif
Publicado: (2026)
Synthetic Data Generation in Low-Resource Settings via Fine-Tuning of Large Language Models
por: Kaddour, Jean, et al.
Publicado: (2023)
por: Kaddour, Jean, et al.
Publicado: (2023)
On Limitations of LLM as Annotator for Low Resource Languages
por: Jadhav, Suramya, et al.
Publicado: (2024)
por: Jadhav, Suramya, et al.
Publicado: (2024)
Aligning Large Language Models to Low-Resource Languages through LLM-Based Selective Translation: A Systematic Study
por: Paul, Rakesh, et al.
Publicado: (2025)
por: Paul, Rakesh, et al.
Publicado: (2025)
L3Cube-IndicHeadline-ID: A Dataset for Headline Identification and Semantic Evaluation in Low-Resource Indian Languages
por: Tanksale, Nishant, et al.
Publicado: (2025)
por: Tanksale, Nishant, et al.
Publicado: (2025)
UPRPRC: Unified Pipeline for Reproducing Parallel Resources -- Corpus from the United Nations
por: Lu, Qiuyang, et al.
Publicado: (2025)
por: Lu, Qiuyang, et al.
Publicado: (2025)
MURI: High-Quality Instruction Tuning Datasets for Low-Resource Languages via Reverse Instructions
por: Köksal, Abdullatif, et al.
Publicado: (2024)
por: Köksal, Abdullatif, et al.
Publicado: (2024)
Comparative Analysis of Different Efficient Fine Tuning Methods of Large Language Models (LLMs) in Low-Resource Setting
por: Srinivasan, Krishna Prasad Varadarajan, et al.
Publicado: (2024)
por: Srinivasan, Krishna Prasad Varadarajan, et al.
Publicado: (2024)
Bridging the Gap: Enhancing LLM Performance for Low-Resource African Languages with New Benchmarks, Fine-Tuning, and Cultural Adjustments
por: Alhanai, Tuka, et al.
Publicado: (2024)
por: Alhanai, Tuka, et al.
Publicado: (2024)
Grammatical Error Correction for Low-Resource Languages: The Case of Zarma
por: Keita, Mamadou K., et al.
Publicado: (2024)
por: Keita, Mamadou K., et al.
Publicado: (2024)
Few-Shot Cross-Lingual Transfer for Prompting Large Language Models in Low-Resource Languages
por: Toukmaji, Christopher
Publicado: (2024)
por: Toukmaji, Christopher
Publicado: (2024)
Evaluating Fine-Tuned LLM Model For Medical Transcription With Small Low-Resource Languages Validated Dataset
por: Chowdhury, Mohammed Nowshad Ruhani, et al.
Publicado: (2026)
por: Chowdhury, Mohammed Nowshad Ruhani, et al.
Publicado: (2026)
Towards Enhancing Health Coaching Dialogue in Low-Resource Settings
por: Zhou, Yue, et al.
Publicado: (2024)
por: Zhou, Yue, et al.
Publicado: (2024)
A Large-Scale Vision-Language Dataset Derived from Open Scientific Literature to Advance Biomedical Generalist AI
por: Lozano, Alejandro, et al.
Publicado: (2025)
por: Lozano, Alejandro, et al.
Publicado: (2025)
ResumeAtlas: Revisiting Resume Classification with Large-Scale Datasets and Large Language Models
por: Heakl, Ahmed, et al.
Publicado: (2024)
por: Heakl, Ahmed, et al.
Publicado: (2024)
On Importance of Layer Pruning for Smaller BERT Models and Low Resource Languages
por: Shirke, Mayur, et al.
Publicado: (2025)
por: Shirke, Mayur, et al.
Publicado: (2025)
Transformer-Driven Triple Fusion Framework for Enhanced Multimodal Author Intent Classification in Low-Resource Bangla
por: Islam, Ariful, et al.
Publicado: (2025)
por: Islam, Ariful, et al.
Publicado: (2025)
KenSwQuAD -- A Question Answering Dataset for Swahili Low Resource Language
por: Wanjawa, Barack W., et al.
Publicado: (2022)
por: Wanjawa, Barack W., et al.
Publicado: (2022)
Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training
por: Kesgin, H. Toprak, et al.
Publicado: (2024)
por: Kesgin, H. Toprak, et al.
Publicado: (2024)
Lugha-Llama: Adapting Large Language Models for African Languages
por: Buzaaba, Happy, et al.
Publicado: (2025)
por: Buzaaba, Happy, et al.
Publicado: (2025)
ForeCite: Adapting Pre-Trained Language Models to Predict Future Citation Rates of Academic Papers
por: Hull, Gavin, et al.
Publicado: (2025)
por: Hull, Gavin, et al.
Publicado: (2025)
LakotaBERT: A Transformer-based Model for Low Resource Lakota Language
por: Parankusham, Kanishka, et al.
Publicado: (2025)
por: Parankusham, Kanishka, et al.
Publicado: (2025)
A Survey on Automatic Online Hate Speech Detection in Low-Resource Languages
por: Das, Susmita, et al.
Publicado: (2024)
por: Das, Susmita, et al.
Publicado: (2024)
Predicting Machine Translation Performance on Low-Resource Languages: The Role of Domain Similarity
por: Khiu, Eric, et al.
Publicado: (2024)
por: Khiu, Eric, et al.
Publicado: (2024)
Generalizing Large Language Model Usability Across Resource-Constrained
por: Tsai, Yun-Da
Publicado: (2025)
por: Tsai, Yun-Da
Publicado: (2025)
Improving Low-Resource Machine Translation via Cross-Linguistic Transfer from Typologically Similar High-Resource Languages
por: Boujkian, Saughmon
Publicado: (2024)
por: Boujkian, Saughmon
Publicado: (2024)
RedPajama: an Open Dataset for Training Large Language Models
por: Weber, Maurice, et al.
Publicado: (2024)
por: Weber, Maurice, et al.
Publicado: (2024)
Large Language Model Prompt Datasets: An In-depth Analysis and Insights
por: Zhang, Yuanming, et al.
Publicado: (2025)
por: Zhang, Yuanming, et al.
Publicado: (2025)
A Multilingual Sentiment Lexicon for Low-Resource Language Translation using Large Languages Models and Explainable AI
por: Malinga, Melusi, et al.
Publicado: (2024)
por: Malinga, Melusi, et al.
Publicado: (2024)
Ejemplares similares
-
Zero-Shot Morphological Discovery in Low-Resource Bantu Languages via Cross-Lingual Transfer and Unsupervised Clustering
por: Mutisya, Hillary, et al.
Publicado: (2026) -
Neural Recovery of Historical Lexical Structure in Bantu Languages from Modern Data
por: Mutisya, Hillary, et al.
Publicado: (2026) -
Attention Sinks in Massively Multilingual Neural Machine Translation:Discovery, Analysis, and Mitigation
por: Mutisya, Hillary, et al.
Publicado: (2026) -
Continued Pretraining for Low-Resource Swahili ASR: Achieving State-of-the-Art Performance with Minimal Labeled Data
por: Mutisya, Hillary, et al.
Publicado: (2026) -
A Hybrid Protocol for Large-Scale Semantic Dataset Generation in Low-Resource Languages: The Turkish Semantic Relations Corpus
por: Tosun, Ebubekir, et al.
Publicado: (2026)