Nanbeige4-3B Technical Report: Exploring the Frontier of Small Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866917129473753088 |
|---|---|
| author | Yang, Chen Peng, Guangyue Zhu, Jiaying Le, Ran Feng, Ruixiang Zhang, Tao Ruan, Wei Liu, Xiaoqi Cheng, Xiaoxue Xu, Xiyun Song, Yang Gao, Yanzipeng Jia, Yiming Xing, Yun Wen, Yuntao Wang, Zekai An, Zhenwei Sun, Zhicong Chen, Zongchao |
| author_facet | Yang, Chen Peng, Guangyue Zhu, Jiaying Le, Ran Feng, Ruixiang Zhang, Tao Ruan, Wei Liu, Xiaoqi Cheng, Xiaoxue Xu, Xiyun Song, Yang Gao, Yanzipeng Jia, Yiming Xing, Yun Wen, Yuntao Wang, Zekai An, Zhenwei Sun, Zhicong Chen, Zongchao |
| contents | We present Nanbeige4-3B, a family of small-scale but high-performing language models. Pretrained on 23T high-quality tokens and finetuned on over 30 million diverse instructions, we extend the boundary of the scaling law for small language models. In pre-training, we design a Fine-Grained Warmup-Stable-Decay (FG-WSD) training scheduler, which progressively refines data mixtures across stages to boost model performance. In post-training, to improve the quality of the SFT data, we design a joint mechanism that integrates deliberative generation refinement and chain-of-thought reconstruction, yielding substantial gains on complex tasks. Following SFT, we employ our flagship reasoning model to distill Nanbeige4-3B through our proposed Dual Preference Distillation (DPD) method, which leads to further performance gains. Finally, a multi-stage reinforcement learning phase was applied, leveraging verifiable rewards and preference modeling to strengthen abilities on both reasoning and human alignment. Extensive evaluations show that Nanbeige4-3B not only significantly outperforms models of comparable parameter scale but also rivals much larger models across a wide range of benchmarks. The model checkpoints are available at https://huggingface.co/Nanbeige. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_06266 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Nanbeige4-3B Technical Report: Exploring the Frontier of Small Language Models Yang, Chen Peng, Guangyue Zhu, Jiaying Le, Ran Feng, Ruixiang Zhang, Tao Ruan, Wei Liu, Xiaoqi Cheng, Xiaoxue Xu, Xiyun Song, Yang Gao, Yanzipeng Jia, Yiming Xing, Yun Wen, Yuntao Wang, Zekai An, Zhenwei Sun, Zhicong Chen, Zongchao Computation and Language We present Nanbeige4-3B, a family of small-scale but high-performing language models. Pretrained on 23T high-quality tokens and finetuned on over 30 million diverse instructions, we extend the boundary of the scaling law for small language models. In pre-training, we design a Fine-Grained Warmup-Stable-Decay (FG-WSD) training scheduler, which progressively refines data mixtures across stages to boost model performance. In post-training, to improve the quality of the SFT data, we design a joint mechanism that integrates deliberative generation refinement and chain-of-thought reconstruction, yielding substantial gains on complex tasks. Following SFT, we employ our flagship reasoning model to distill Nanbeige4-3B through our proposed Dual Preference Distillation (DPD) method, which leads to further performance gains. Finally, a multi-stage reinforcement learning phase was applied, leveraging verifiable rewards and preference modeling to strengthen abilities on both reasoning and human alignment. Extensive evaluations show that Nanbeige4-3B not only significantly outperforms models of comparable parameter scale but also rivals much larger models across a wide range of benchmarks. The model checkpoints are available at https://huggingface.co/Nanbeige. |
| title | Nanbeige4-3B Technical Report: Exploring the Frontier of Small Language Models |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2512.06266 |