Nanbeige4.1-3B: A Small General Model that Reasons, Aligns, and Acts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Chen, Peng, Guangyue, Zhu, Jiaying, Le, Ran, Feng, Ruixiang, Zhang, Tao, Xu, Xiyun, Song, Yang, Jia, Yiming, Wen, Yuntao, Xu, Yunzhi, Wang, Zekai, An, Zhenwei, Sun, Zhicong, Chen, Zongchao
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918337760460800
author Yang, Chen
Peng, Guangyue
Zhu, Jiaying
Le, Ran
Feng, Ruixiang
Zhang, Tao
Xu, Xiyun
Song, Yang
Jia, Yiming
Wen, Yuntao
Xu, Yunzhi
Wang, Zekai
An, Zhenwei
Sun, Zhicong
Chen, Zongchao
author_facet Yang, Chen
Peng, Guangyue
Zhu, Jiaying
Le, Ran
Feng, Ruixiang
Zhang, Tao
Xu, Xiyun
Song, Yang
Jia, Yiming
Wen, Yuntao
Xu, Yunzhi
Wang, Zekai
An, Zhenwei
Sun, Zhicong
Chen, Zongchao
contents We present Nanbeige4.1-3B, a unified generalist language model that simultaneously achieves strong agentic behavior, code generation, and general reasoning with only 3B parameters. To the best of our knowledge, it is the first open-source small language model (SLM) to achieve such versatility in a single model. To improve reasoning and preference alignment, we combine point-wise and pair-wise reward modeling, ensuring high-quality, human-aligned responses. For code generation, we design complexity-aware rewards in Reinforcement Learning, optimizing both correctness and efficiency. In deep search, we perform complex data synthesis and incorporate turn-level supervision during training. This enables stable long-horizon tool interactions, allowing Nanbeige4.1-3B to reliably execute up to 600 tool-call turns for complex problem-solving. Extensive experimental results show that Nanbeige4.1-3B significantly outperforms prior models of similar scale, such as Nanbeige4-3B-2511 and Qwen3-4B, even achieving superior performance compared to much larger models, such as Qwen3-30B-A3B. Our results demonstrate that small models can achieve both broad competence and strong specialization simultaneously, redefining the potential of 3B parameter models.
format Preprint
id arxiv_https___arxiv_org_abs_2602_13367
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Nanbeige4.1-3B: A Small General Model that Reasons, Aligns, and Acts
Yang, Chen
Peng, Guangyue
Zhu, Jiaying
Le, Ran
Feng, Ruixiang
Zhang, Tao
Xu, Xiyun
Song, Yang
Jia, Yiming
Wen, Yuntao
Xu, Yunzhi
Wang, Zekai
An, Zhenwei
Sun, Zhicong
Chen, Zongchao
Artificial Intelligence
Computation and Language
We present Nanbeige4.1-3B, a unified generalist language model that simultaneously achieves strong agentic behavior, code generation, and general reasoning with only 3B parameters. To the best of our knowledge, it is the first open-source small language model (SLM) to achieve such versatility in a single model. To improve reasoning and preference alignment, we combine point-wise and pair-wise reward modeling, ensuring high-quality, human-aligned responses. For code generation, we design complexity-aware rewards in Reinforcement Learning, optimizing both correctness and efficiency. In deep search, we perform complex data synthesis and incorporate turn-level supervision during training. This enables stable long-horizon tool interactions, allowing Nanbeige4.1-3B to reliably execute up to 600 tool-call turns for complex problem-solving. Extensive experimental results show that Nanbeige4.1-3B significantly outperforms prior models of similar scale, such as Nanbeige4-3B-2511 and Qwen3-4B, even achieving superior performance compared to much larger models, such as Qwen3-30B-A3B. Our results demonstrate that small models can achieve both broad competence and strong specialization simultaneously, redefining the potential of 3B parameter models.
title Nanbeige4.1-3B: A Small General Model that Reasons, Aligns, and Acts
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2602.13367