Bielik 11B v3: Multilingual Large Language Model for European Languages

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ociepa, Krzysztof, Flis, Łukasz, Kinas, Remigiusz, Wróbel, Krzysztof, Gwoździej, Adrian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918293457076224
author Ociepa, Krzysztof
Flis, Łukasz
Kinas, Remigiusz
Wróbel, Krzysztof
Gwoździej, Adrian
author_facet Ociepa, Krzysztof
Flis, Łukasz
Kinas, Remigiusz
Wróbel, Krzysztof
Gwoździej, Adrian
contents We present Bielik 11B v3, a state-of-the-art language model highly optimized for the Polish language, while also maintaining strong capabilities in other European languages. This model extends the Mistral 7B v0.2 architecture, scaled to 11B parameters via depth up-scaling. Its development involved a comprehensive four-stage training pipeline: continuous pre-training, supervised fine-tuning (SFT), Direct Preference Optimization (DPO), and reinforcement learning. Comprehensive evaluations demonstrate that Bielik 11B v3 achieves exceptional performance. It significantly surpasses other specialized Polish language models and outperforms many larger models (with 2-6 times more parameters) on a wide range of tasks, from basic linguistic understanding to complex reasoning. The model's parameter efficiency, combined with extensive quantization options, allows for effective deployment across diverse hardware configurations. Bielik 11B v3 not only advances AI capabilities for the Polish language but also establishes a new benchmark for developing resource-efficient, high-performance models for less-represented languages.
format Preprint
id arxiv_https___arxiv_org_abs_2601_11579
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Bielik 11B v3: Multilingual Large Language Model for European Languages
Ociepa, Krzysztof
Flis, Łukasz
Kinas, Remigiusz
Wróbel, Krzysztof
Gwoździej, Adrian
Computation and Language
Artificial Intelligence
I.2.7
We present Bielik 11B v3, a state-of-the-art language model highly optimized for the Polish language, while also maintaining strong capabilities in other European languages. This model extends the Mistral 7B v0.2 architecture, scaled to 11B parameters via depth up-scaling. Its development involved a comprehensive four-stage training pipeline: continuous pre-training, supervised fine-tuning (SFT), Direct Preference Optimization (DPO), and reinforcement learning. Comprehensive evaluations demonstrate that Bielik 11B v3 achieves exceptional performance. It significantly surpasses other specialized Polish language models and outperforms many larger models (with 2-6 times more parameters) on a wide range of tasks, from basic linguistic understanding to complex reasoning. The model's parameter efficiency, combined with extensive quantization options, allows for effective deployment across diverse hardware configurations. Bielik 11B v3 not only advances AI capabilities for the Polish language but also establishes a new benchmark for developing resource-efficient, high-performance models for less-represented languages.
title Bielik 11B v3: Multilingual Large Language Model for European Languages
topic Computation and Language
Artificial Intelligence
I.2.7
url https://arxiv.org/abs/2601.11579