MiniLingua: A Small Open-Source LLM for European Languages

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Aksenova, Anna, Zverkov, Boris, Dainese, Nicola, Nikitin, Alexander, Marttinen, Pekka
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917146971340800
author Aksenova, Anna
Zverkov, Boris
Dainese, Nicola
Nikitin, Alexander
Marttinen, Pekka
author_facet Aksenova, Anna
Zverkov, Boris
Dainese, Nicola
Nikitin, Alexander
Marttinen, Pekka
contents Large language models are powerful but often limited by high computational cost, privacy concerns, and English-centric training. Recent progress demonstrates that small, efficient models with around one billion parameters can deliver strong results and enable on-device use. This paper introduces MiniLingua, a multilingual open-source LLM of one billion parameters trained from scratch for 13 European languages, designed to balance coverage and instruction-following capabilities. Based on evaluation results, the instruction-tuned version of MiniLingua outperforms EuroLLM, a model with a similar training approach but a larger training budget, on summarization, classification and both open- and closed-book question answering. Moreover, it remains competitive with more advanced state-of-the-art models on open-ended generation tasks. We release model weights, tokenizer and source code used for data processing and model training.
format Preprint
id arxiv_https___arxiv_org_abs_2512_13298
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MiniLingua: A Small Open-Source LLM for European Languages
Aksenova, Anna
Zverkov, Boris
Dainese, Nicola
Nikitin, Alexander
Marttinen, Pekka
Computation and Language
Artificial Intelligence
Large language models are powerful but often limited by high computational cost, privacy concerns, and English-centric training. Recent progress demonstrates that small, efficient models with around one billion parameters can deliver strong results and enable on-device use. This paper introduces MiniLingua, a multilingual open-source LLM of one billion parameters trained from scratch for 13 European languages, designed to balance coverage and instruction-following capabilities. Based on evaluation results, the instruction-tuned version of MiniLingua outperforms EuroLLM, a model with a similar training approach but a larger training budget, on summarization, classification and both open- and closed-book question answering. Moreover, it remains competitive with more advanced state-of-the-art models on open-ended generation tasks. We release model weights, tokenizer and source code used for data processing and model training.
title MiniLingua: A Small Open-Source LLM for European Languages
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2512.13298