Nanbeige4-3B Technical Report: Exploring the Frontier of Small Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yang, Chen, Peng, Guangyue, Zhu, Jiaying, Le, Ran, Feng, Ruixiang, Zhang, Tao, Ruan, Wei, Liu, Xiaoqi, Cheng, Xiaoxue, Xu, Xiyun, Song, Yang, Gao, Yanzipeng, Jia, Yiming, Xing, Yun, Wen, Yuntao, Wang, Zekai, An, Zhenwei, Sun, Zhicong, Chen, Zongchao
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917129473753088
author Yang, Chen
Peng, Guangyue
Zhu, Jiaying
Le, Ran
Feng, Ruixiang
Zhang, Tao
Ruan, Wei
Liu, Xiaoqi
Cheng, Xiaoxue
Xu, Xiyun
Song, Yang
Gao, Yanzipeng
Jia, Yiming
Xing, Yun
Wen, Yuntao
Wang, Zekai
An, Zhenwei
Sun, Zhicong
Chen, Zongchao
author_facet Yang, Chen
Peng, Guangyue
Zhu, Jiaying
Le, Ran
Feng, Ruixiang
Zhang, Tao
Ruan, Wei
Liu, Xiaoqi
Cheng, Xiaoxue
Xu, Xiyun
Song, Yang
Gao, Yanzipeng
Jia, Yiming
Xing, Yun
Wen, Yuntao
Wang, Zekai
An, Zhenwei
Sun, Zhicong
Chen, Zongchao
contents We present Nanbeige4-3B, a family of small-scale but high-performing language models. Pretrained on 23T high-quality tokens and finetuned on over 30 million diverse instructions, we extend the boundary of the scaling law for small language models. In pre-training, we design a Fine-Grained Warmup-Stable-Decay (FG-WSD) training scheduler, which progressively refines data mixtures across stages to boost model performance. In post-training, to improve the quality of the SFT data, we design a joint mechanism that integrates deliberative generation refinement and chain-of-thought reconstruction, yielding substantial gains on complex tasks. Following SFT, we employ our flagship reasoning model to distill Nanbeige4-3B through our proposed Dual Preference Distillation (DPD) method, which leads to further performance gains. Finally, a multi-stage reinforcement learning phase was applied, leveraging verifiable rewards and preference modeling to strengthen abilities on both reasoning and human alignment. Extensive evaluations show that Nanbeige4-3B not only significantly outperforms models of comparable parameter scale but also rivals much larger models across a wide range of benchmarks. The model checkpoints are available at https://huggingface.co/Nanbeige.
format Preprint
id arxiv_https___arxiv_org_abs_2512_06266
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Nanbeige4-3B Technical Report: Exploring the Frontier of Small Language Models
Yang, Chen
Peng, Guangyue
Zhu, Jiaying
Le, Ran
Feng, Ruixiang
Zhang, Tao
Ruan, Wei
Liu, Xiaoqi
Cheng, Xiaoxue
Xu, Xiyun
Song, Yang
Gao, Yanzipeng
Jia, Yiming
Xing, Yun
Wen, Yuntao
Wang, Zekai
An, Zhenwei
Sun, Zhicong
Chen, Zongchao
Computation and Language
We present Nanbeige4-3B, a family of small-scale but high-performing language models. Pretrained on 23T high-quality tokens and finetuned on over 30 million diverse instructions, we extend the boundary of the scaling law for small language models. In pre-training, we design a Fine-Grained Warmup-Stable-Decay (FG-WSD) training scheduler, which progressively refines data mixtures across stages to boost model performance. In post-training, to improve the quality of the SFT data, we design a joint mechanism that integrates deliberative generation refinement and chain-of-thought reconstruction, yielding substantial gains on complex tasks. Following SFT, we employ our flagship reasoning model to distill Nanbeige4-3B through our proposed Dual Preference Distillation (DPD) method, which leads to further performance gains. Finally, a multi-stage reinforcement learning phase was applied, leveraging verifiable rewards and preference modeling to strengthen abilities on both reasoning and human alignment. Extensive evaluations show that Nanbeige4-3B not only significantly outperforms models of comparable parameter scale but also rivals much larger models across a wide range of benchmarks. The model checkpoints are available at https://huggingface.co/Nanbeige.
title Nanbeige4-3B Technical Report: Exploring the Frontier of Small Language Models
topic Computation and Language
url https://arxiv.org/abs/2512.06266