InstructLR: A Scalable Approach to Create Instruction Dataset for Under-Resourced Languages

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Keita, Mamadou K., Diarra, Sebastien, Homan, Christopher, Diallo, Seydou
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914308635492352
author Keita, Mamadou K.
Diarra, Sebastien
Homan, Christopher
Diallo, Seydou
author_facet Keita, Mamadou K.
Diarra, Sebastien
Homan, Christopher
Diallo, Seydou
contents Effective text generation and chat interfaces for low-resource languages (LRLs) remain a challenge for state-of-the-art large language models (LLMs) to support. This is mainly due to the difficulty of curating high-quality instruction datasets for LRLs, a limitation prevalent in the languages spoken across the African continent and other regions. Current approaches, such as automated translation and synthetic data generation, frequently yield outputs that lack fluency or even orthographic consistency. In this paper, we introduce InstructLR, a novel framework designed to generate high-quality instruction datasets for LRLs. Our approach integrates LLM-driven text generation with a dual-layer quality filtering mechanism: an automated filtering layer based on retrieval-augmented-generation (RAG)-based n-shot prompting, and a human-in-the-loop validation layer. Drawing inspiration from benchmarks such as MMLU in task definition, InstructLR has facilitated the creation of three multi-domain instruction benchmarks: ZarmaInstruct-50k, BambaraInstruct-50k, and FulfuldeInstruct-50k.
format Preprint
id arxiv_https___arxiv_org_abs_2512_02213
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle InstructLR: A Scalable Approach to Create Instruction Dataset for Under-Resourced Languages
Keita, Mamadou K.
Diarra, Sebastien
Homan, Christopher
Diallo, Seydou
Machine Learning
Effective text generation and chat interfaces for low-resource languages (LRLs) remain a challenge for state-of-the-art large language models (LLMs) to support. This is mainly due to the difficulty of curating high-quality instruction datasets for LRLs, a limitation prevalent in the languages spoken across the African continent and other regions. Current approaches, such as automated translation and synthetic data generation, frequently yield outputs that lack fluency or even orthographic consistency. In this paper, we introduce InstructLR, a novel framework designed to generate high-quality instruction datasets for LRLs. Our approach integrates LLM-driven text generation with a dual-layer quality filtering mechanism: an automated filtering layer based on retrieval-augmented-generation (RAG)-based n-shot prompting, and a human-in-the-loop validation layer. Drawing inspiration from benchmarks such as MMLU in task definition, InstructLR has facilitated the creation of three multi-domain instruction benchmarks: ZarmaInstruct-50k, BambaraInstruct-50k, and FulfuldeInstruct-50k.
title InstructLR: A Scalable Approach to Create Instruction Dataset for Under-Resourced Languages
topic Machine Learning
url https://arxiv.org/abs/2512.02213