ConLID: Supervised Contrastive Learning for Low-Resource Language Identification

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Foroutan, Negar, Saydaliev, Jakhongir, Kim, Ye Eun, Bosselut, Antoine
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917326482309120
author Foroutan, Negar
Saydaliev, Jakhongir
Kim, Ye Eun
Bosselut, Antoine
author_facet Foroutan, Negar
Saydaliev, Jakhongir
Kim, Ye Eun
Bosselut, Antoine
contents Language identification (LID) is a critical step in curating multilingual LLM pretraining corpora from web crawls. While many studies on LID model training focus on collecting diverse training data to improve performance, low-resource languages -- often limited to single-domain data, such as the Bible -- continue to perform poorly. To resolve these imbalance and bias issues, we propose a novel supervised contrastive learning (SCL) approach to learn domain-invariant representations for low-resource languages. We show that our approach improves LID performance on out-of-domain data for low-resource languages by 3.2 percentage points, while maintaining its performance for the high-resource languages.
format Preprint
id arxiv_https___arxiv_org_abs_2506_15304
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ConLID: Supervised Contrastive Learning for Low-Resource Language Identification
Foroutan, Negar
Saydaliev, Jakhongir
Kim, Ye Eun
Bosselut, Antoine
Computation and Language
Artificial Intelligence
Machine Learning
Language identification (LID) is a critical step in curating multilingual LLM pretraining corpora from web crawls. While many studies on LID model training focus on collecting diverse training data to improve performance, low-resource languages -- often limited to single-domain data, such as the Bible -- continue to perform poorly. To resolve these imbalance and bias issues, we propose a novel supervised contrastive learning (SCL) approach to learn domain-invariant representations for low-resource languages. We show that our approach improves LID performance on out-of-domain data for low-resource languages by 3.2 percentage points, while maintaining its performance for the high-resource languages.
title ConLID: Supervised Contrastive Learning for Low-Resource Language Identification
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2506.15304