TildeOpen LLM: Leveraging Curriculum Learning to Achieve Equitable Language Representation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bergmanis, Toms, Kronis, Martins, Pretkalniņš, Ingus Jānis, Nicmanis, Dāvis, Jelinska, Jeļizaveta, Rozis, Roberts, Vīksna, Rinalds, Pinnis, Mārcis
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917446554746880
author Bergmanis, Toms
Kronis, Martins
Pretkalniņš, Ingus Jānis
Nicmanis, Dāvis
Jelinska, Jeļizaveta
Rozis, Roberts
Vīksna, Rinalds
Pinnis, Mārcis
author_facet Bergmanis, Toms
Kronis, Martins
Pretkalniņš, Ingus Jānis
Nicmanis, Dāvis
Jelinska, Jeļizaveta
Rozis, Roberts
Vīksna, Rinalds
Pinnis, Mārcis
contents Large language models often underperform in many European languages due to the dominance of English and a few high-resource languages in training data. This paper presents TildeOpen LLM, a 30-billion-parameter open-weight foundational model trained for 34 European languages to promote linguistic equity and improve performance for low-resource languages. To address the data imbalance, we combine dataset upsampling with a curriculum-based training schedule that alternates between uniform and natural language distributions. The resulting model performs favorably compared to other multilingual LLMs despite being trained with significantly fewer computing resources. Evaluation across multiple multilingual benchmarks shows that TildeOpen surpasses existing open-weight models in text generation and comprehension, particularly for Baltic, Finno-Ugric, and Slavic languages. Human evaluations confirm an up to tenfold reduction in linguistic errors relative to leading baselines. The model and associated resources are fully open-weight and publicly available at huggingface.co/TildeAI/TildeOpen-30b. These outcomes demonstrate that careful data curation and balanced training strategies can substantially enhance multilingual model quality without increasing model size or training volume.
format Preprint
id arxiv_https___arxiv_org_abs_2603_08182
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle TildeOpen LLM: Leveraging Curriculum Learning to Achieve Equitable Language Representation
Bergmanis, Toms
Kronis, Martins
Pretkalniņš, Ingus Jānis
Nicmanis, Dāvis
Jelinska, Jeļizaveta
Rozis, Roberts
Vīksna, Rinalds
Pinnis, Mārcis
Computation and Language
Artificial Intelligence
Large language models often underperform in many European languages due to the dominance of English and a few high-resource languages in training data. This paper presents TildeOpen LLM, a 30-billion-parameter open-weight foundational model trained for 34 European languages to promote linguistic equity and improve performance for low-resource languages. To address the data imbalance, we combine dataset upsampling with a curriculum-based training schedule that alternates between uniform and natural language distributions. The resulting model performs favorably compared to other multilingual LLMs despite being trained with significantly fewer computing resources. Evaluation across multiple multilingual benchmarks shows that TildeOpen surpasses existing open-weight models in text generation and comprehension, particularly for Baltic, Finno-Ugric, and Slavic languages. Human evaluations confirm an up to tenfold reduction in linguistic errors relative to leading baselines. The model and associated resources are fully open-weight and publicly available at huggingface.co/TildeAI/TildeOpen-30b. These outcomes demonstrate that careful data curation and balanced training strategies can substantially enhance multilingual model quality without increasing model size or training volume.
title TildeOpen LLM: Leveraging Curriculum Learning to Achieve Equitable Language Representation
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2603.08182