Gamayun's Path to Multilingual Mastery: Cost-Efficient Training of a 1.5B-Parameter LLM

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Podolskiy, Alexander, Molokov, Semen, Gerasin, Timofey, Titov, Maksim, Rukhovich, Alexey, Khrapov, Artem, Morozov, Kirill, Tetin, Evgeny, Korikov, Constantine, Efimov, Pavel, Lazukova, Polina, Skripkar, Yuliya, Okhotnikov, Nikita, Piontkovskaya, Irina, Xiaojun, Meng, Xueyi, Zou, Zhenhe, Zhang
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909977393758208
author Podolskiy, Alexander
Molokov, Semen
Gerasin, Timofey
Titov, Maksim
Rukhovich, Alexey
Khrapov, Artem
Morozov, Kirill
Tetin, Evgeny
Korikov, Constantine
Efimov, Pavel
Lazukova, Polina
Skripkar, Yuliya
Okhotnikov, Nikita
Piontkovskaya, Irina
Xiaojun, Meng
Xueyi, Zou
Zhenhe, Zhang
author_facet Podolskiy, Alexander
Molokov, Semen
Gerasin, Timofey
Titov, Maksim
Rukhovich, Alexey
Khrapov, Artem
Morozov, Kirill
Tetin, Evgeny
Korikov, Constantine
Efimov, Pavel
Lazukova, Polina
Skripkar, Yuliya
Okhotnikov, Nikita
Piontkovskaya, Irina
Xiaojun, Meng
Xueyi, Zou
Zhenhe, Zhang
contents We present Gamayun, a 1.5B-parameter multilingual language model trained entirely from scratch on 2.5T tokens. Designed for efficiency and deployment in resource-constrained environments, Gamayun addresses the lack of research on small non-English-centric LLMs by adopting a novel two-stage pre-training strategy: balanced multilingual training for cross-lingual alignment, followed by high-quality English enrichment to transfer performance gains across languages. Our model supports 12 languages, with special focus on Russian. Despite a significantly smaller training budget than comparable models, Gamayun outperforms LLaMA3.2-1B (9T tokens) on all considered benchmarks, and surpasses Qwen2.5-1.5B (18T tokens) on a wide range of English and multilingual tasks. It matches or exceeds Qwen3 (36T tokens) on most tasks outside advanced STEM, achieving state-of-the-art results in Russian, including the MERA benchmark, among the models of comparable size (1-2B parameters).
format Preprint
id arxiv_https___arxiv_org_abs_2512_21580
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Gamayun's Path to Multilingual Mastery: Cost-Efficient Training of a 1.5B-Parameter LLM
Podolskiy, Alexander
Molokov, Semen
Gerasin, Timofey
Titov, Maksim
Rukhovich, Alexey
Khrapov, Artem
Morozov, Kirill
Tetin, Evgeny
Korikov, Constantine
Efimov, Pavel
Lazukova, Polina
Skripkar, Yuliya
Okhotnikov, Nikita
Piontkovskaya, Irina
Xiaojun, Meng
Xueyi, Zou
Zhenhe, Zhang
Computation and Language
We present Gamayun, a 1.5B-parameter multilingual language model trained entirely from scratch on 2.5T tokens. Designed for efficiency and deployment in resource-constrained environments, Gamayun addresses the lack of research on small non-English-centric LLMs by adopting a novel two-stage pre-training strategy: balanced multilingual training for cross-lingual alignment, followed by high-quality English enrichment to transfer performance gains across languages. Our model supports 12 languages, with special focus on Russian. Despite a significantly smaller training budget than comparable models, Gamayun outperforms LLaMA3.2-1B (9T tokens) on all considered benchmarks, and surpasses Qwen2.5-1.5B (18T tokens) on a wide range of English and multilingual tasks. It matches or exceeds Qwen3 (36T tokens) on most tasks outside advanced STEM, achieving state-of-the-art results in Russian, including the MERA benchmark, among the models of comparable size (1-2B parameters).
title Gamayun's Path to Multilingual Mastery: Cost-Efficient Training of a 1.5B-Parameter LLM
topic Computation and Language
url https://arxiv.org/abs/2512.21580