ConLID: Supervised Contrastive Learning for Low-Resource Language Identification
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866917326482309120 |
|---|---|
| author | Foroutan, Negar Saydaliev, Jakhongir Kim, Ye Eun Bosselut, Antoine |
| author_facet | Foroutan, Negar Saydaliev, Jakhongir Kim, Ye Eun Bosselut, Antoine |
| contents | Language identification (LID) is a critical step in curating multilingual LLM pretraining corpora from web crawls. While many studies on LID model training focus on collecting diverse training data to improve performance, low-resource languages -- often limited to single-domain data, such as the Bible -- continue to perform poorly. To resolve these imbalance and bias issues, we propose a novel supervised contrastive learning (SCL) approach to learn domain-invariant representations for low-resource languages. We show that our approach improves LID performance on out-of-domain data for low-resource languages by 3.2 percentage points, while maintaining its performance for the high-resource languages. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_15304 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | ConLID: Supervised Contrastive Learning for Low-Resource Language Identification Foroutan, Negar Saydaliev, Jakhongir Kim, Ye Eun Bosselut, Antoine Computation and Language Artificial Intelligence Machine Learning Language identification (LID) is a critical step in curating multilingual LLM pretraining corpora from web crawls. While many studies on LID model training focus on collecting diverse training data to improve performance, low-resource languages -- often limited to single-domain data, such as the Bible -- continue to perform poorly. To resolve these imbalance and bias issues, we propose a novel supervised contrastive learning (SCL) approach to learn domain-invariant representations for low-resource languages. We show that our approach improves LID performance on out-of-domain data for low-resource languages by 3.2 percentage points, while maintaining its performance for the high-resource languages. |
| title | ConLID: Supervised Contrastive Learning for Low-Resource Language Identification |
| topic | Computation and Language Artificial Intelligence Machine Learning |
| url | https://arxiv.org/abs/2506.15304 |