Gamayun's Path to Multilingual Mastery: Cost-Efficient Training of a 1.5B-Parameter LLM
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866909977393758208 |
|---|---|
| author | Podolskiy, Alexander Molokov, Semen Gerasin, Timofey Titov, Maksim Rukhovich, Alexey Khrapov, Artem Morozov, Kirill Tetin, Evgeny Korikov, Constantine Efimov, Pavel Lazukova, Polina Skripkar, Yuliya Okhotnikov, Nikita Piontkovskaya, Irina Xiaojun, Meng Xueyi, Zou Zhenhe, Zhang |
| author_facet | Podolskiy, Alexander Molokov, Semen Gerasin, Timofey Titov, Maksim Rukhovich, Alexey Khrapov, Artem Morozov, Kirill Tetin, Evgeny Korikov, Constantine Efimov, Pavel Lazukova, Polina Skripkar, Yuliya Okhotnikov, Nikita Piontkovskaya, Irina Xiaojun, Meng Xueyi, Zou Zhenhe, Zhang |
| contents | We present Gamayun, a 1.5B-parameter multilingual language model trained entirely from scratch on 2.5T tokens. Designed for efficiency and deployment in resource-constrained environments, Gamayun addresses the lack of research on small non-English-centric LLMs by adopting a novel two-stage pre-training strategy: balanced multilingual training for cross-lingual alignment, followed by high-quality English enrichment to transfer performance gains across languages. Our model supports 12 languages, with special focus on Russian. Despite a significantly smaller training budget than comparable models, Gamayun outperforms LLaMA3.2-1B (9T tokens) on all considered benchmarks, and surpasses Qwen2.5-1.5B (18T tokens) on a wide range of English and multilingual tasks. It matches or exceeds Qwen3 (36T tokens) on most tasks outside advanced STEM, achieving state-of-the-art results in Russian, including the MERA benchmark, among the models of comparable size (1-2B parameters). |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_21580 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Gamayun's Path to Multilingual Mastery: Cost-Efficient Training of a 1.5B-Parameter LLM Podolskiy, Alexander Molokov, Semen Gerasin, Timofey Titov, Maksim Rukhovich, Alexey Khrapov, Artem Morozov, Kirill Tetin, Evgeny Korikov, Constantine Efimov, Pavel Lazukova, Polina Skripkar, Yuliya Okhotnikov, Nikita Piontkovskaya, Irina Xiaojun, Meng Xueyi, Zou Zhenhe, Zhang Computation and Language We present Gamayun, a 1.5B-parameter multilingual language model trained entirely from scratch on 2.5T tokens. Designed for efficiency and deployment in resource-constrained environments, Gamayun addresses the lack of research on small non-English-centric LLMs by adopting a novel two-stage pre-training strategy: balanced multilingual training for cross-lingual alignment, followed by high-quality English enrichment to transfer performance gains across languages. Our model supports 12 languages, with special focus on Russian. Despite a significantly smaller training budget than comparable models, Gamayun outperforms LLaMA3.2-1B (9T tokens) on all considered benchmarks, and surpasses Qwen2.5-1.5B (18T tokens) on a wide range of English and multilingual tasks. It matches or exceeds Qwen3 (36T tokens) on most tasks outside advanced STEM, achieving state-of-the-art results in Russian, including the MERA benchmark, among the models of comparable size (1-2B parameters). |
| title | Gamayun's Path to Multilingual Mastery: Cost-Efficient Training of a 1.5B-Parameter LLM |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2512.21580 |