LLMic: Romanian Foundation Language Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bădoiu, Vlad-Andrei, Dumitru, Mihai-Valentin, Gherghescu, Alexandru M., Agache, Alexandru, Raiciu, Costin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917892191158272
author Bădoiu, Vlad-Andrei
Dumitru, Mihai-Valentin
Gherghescu, Alexandru M.
Agache, Alexandru
Raiciu, Costin
author_facet Bădoiu, Vlad-Andrei
Dumitru, Mihai-Valentin
Gherghescu, Alexandru M.
Agache, Alexandru
Raiciu, Costin
contents Recent advances in Large Language Models (LLMs) have demonstrated remarkable capabilities across various tasks with commercial models leading the way. While open models usually operate at a smaller scale, they maintain competitiveness through specialization and fine-tuning. However, a significant challenge persists: open models often underperform in low-resource languages due to limited representation in the training corpus. In this paper, we present LLMic, a bilingual foundation language model designed specifically for the Romanian Language. We document the complete process of pretraining a foundation model for a low-resource language, including corpus construction, architecture selection, and hyper-parameter optimization. Our evaluation demonstrates that LLMic can be specialized for tasks in the target language, achieving results comparable to other much larger open models. We show that fine-tuning LLMic for language translation after the initial pretraining phase outperforms existing solutions in English-to-Romanian translation tasks. This opens the path for efficient large-scale processing for the Romanian language community, using the much smaller LLMic model
format Preprint
id arxiv_https___arxiv_org_abs_2501_07721
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LLMic: Romanian Foundation Language Model
Bădoiu, Vlad-Andrei
Dumitru, Mihai-Valentin
Gherghescu, Alexandru M.
Agache, Alexandru
Raiciu, Costin
Computation and Language
Recent advances in Large Language Models (LLMs) have demonstrated remarkable capabilities across various tasks with commercial models leading the way. While open models usually operate at a smaller scale, they maintain competitiveness through specialization and fine-tuning. However, a significant challenge persists: open models often underperform in low-resource languages due to limited representation in the training corpus. In this paper, we present LLMic, a bilingual foundation language model designed specifically for the Romanian Language. We document the complete process of pretraining a foundation model for a low-resource language, including corpus construction, architecture selection, and hyper-parameter optimization. Our evaluation demonstrates that LLMic can be specialized for tasks in the target language, achieving results comparable to other much larger open models. We show that fine-tuning LLMic for language translation after the initial pretraining phase outperforms existing solutions in English-to-Romanian translation tasks. This opens the path for efficient large-scale processing for the Romanian language community, using the much smaller LLMic model
title LLMic: Romanian Foundation Language Model
topic Computation and Language
url https://arxiv.org/abs/2501.07721