Low-Resource Dialect Adaptation of Large Language Models: A French Dialect Case-Study

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Khan, Eeham, Saidani, Firas, Van Esbroeck, Owen, Khoury, Richard, Kosseim, Leila
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918358869344256
author Khan, Eeham
Saidani, Firas
Van Esbroeck, Owen
Khoury, Richard
Kosseim, Leila
author_facet Khan, Eeham
Saidani, Firas
Van Esbroeck, Owen
Khoury, Richard
Kosseim, Leila
contents Despite the widespread adoption of Large Language Models (LLMs), their strongest capabilities remain largely confined to a small number of high-resource languages for which there is abundant training data. Recently, continual pre-training (CPT) has emerged as a means to fine-tune these models to low-resource regional dialects. In this paper, we study the use of CPT for dialect learning under tight data and compute budgets. Using low-rank adaptation (LoRA) and compute-efficient continual pre-training, we adapt three LLMs to the Québec French dialect using a very small dataset and benchmark them on the COLE suite. Our experiments demonstrate an improvement on the minority dialect benchmarks with minimal regression on the prestige language benchmarks with around 1% of model parameters updated. Analysis of the results demonstrate that gains are highly contingent on corpus composition. These findings indicate that CPT with parameter-efficient fine-tuning (PEFT) can narrow the dialect gap by providing cost-effective and sustainable language resource creation, expanding high-quality LLM access to minority linguistic communities. To support reproducibility and broaden access, we release the first Québec French LLMs on Hugging Face.
format Preprint
id arxiv_https___arxiv_org_abs_2510_22747
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Low-Resource Dialect Adaptation of Large Language Models: A French Dialect Case-Study
Khan, Eeham
Saidani, Firas
Van Esbroeck, Owen
Khoury, Richard
Kosseim, Leila
Computation and Language
Artificial Intelligence
Despite the widespread adoption of Large Language Models (LLMs), their strongest capabilities remain largely confined to a small number of high-resource languages for which there is abundant training data. Recently, continual pre-training (CPT) has emerged as a means to fine-tune these models to low-resource regional dialects. In this paper, we study the use of CPT for dialect learning under tight data and compute budgets. Using low-rank adaptation (LoRA) and compute-efficient continual pre-training, we adapt three LLMs to the Québec French dialect using a very small dataset and benchmark them on the COLE suite. Our experiments demonstrate an improvement on the minority dialect benchmarks with minimal regression on the prestige language benchmarks with around 1% of model parameters updated. Analysis of the results demonstrate that gains are highly contingent on corpus composition. These findings indicate that CPT with parameter-efficient fine-tuning (PEFT) can narrow the dialect gap by providing cost-effective and sustainable language resource creation, expanding high-quality LLM access to minority linguistic communities. To support reproducibility and broaden access, we release the first Québec French LLMs on Hugging Face.
title Low-Resource Dialect Adaptation of Large Language Models: A French Dialect Case-Study
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.22747