Thinking with DistilQwen: A Tale of Four Distilled Reasoning and Reward Model Series

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Cai, Wenrui, Wang, Chengyu, Yan, Junbing, Huang, Jun, Fang, Xiangzhong
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911247070396416
author Cai, Wenrui
Wang, Chengyu
Yan, Junbing
Huang, Jun
Fang, Xiangzhong
author_facet Cai, Wenrui
Wang, Chengyu
Yan, Junbing
Huang, Jun
Fang, Xiangzhong
contents Recently, the demand for small and efficient reasoning models to support real-world applications has driven the development of knowledge distillation techniques that balance reasoning performance and inference speed. In this paper, we further extend the DistilQwen model family, initialized from the Qwen models, by introducing four model series specifically designed to meet industrial requirements. The distilled model collection comprises: (1) slow-thinking models, optimized for reasoning tasks that require high accuracy; (2) two series of adaptive-thinking models, which dynamically adjust reasoning strategies based on input tasks to maximize efficiency across diverse scenarios; and (3) distilled reward models, which enable further reinforcement learning of reasoning models using distilled knowledge. Comprehensive evaluations across multiple benchmarks demonstrate both high inference efficiency and strong reasoning performance for these models, as well as the practical utility of distilled reward models. We further show that these models support industry practitioners by providing scalable training and inference functionalities on the Alibaba Cloud PAI (Platform for Artificial Intelligence) platform.
format Preprint
id arxiv_https___arxiv_org_abs_2511_01354
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Thinking with DistilQwen: A Tale of Four Distilled Reasoning and Reward Model Series
Cai, Wenrui
Wang, Chengyu
Yan, Junbing
Huang, Jun
Fang, Xiangzhong
Computation and Language
Artificial Intelligence
Recently, the demand for small and efficient reasoning models to support real-world applications has driven the development of knowledge distillation techniques that balance reasoning performance and inference speed. In this paper, we further extend the DistilQwen model family, initialized from the Qwen models, by introducing four model series specifically designed to meet industrial requirements. The distilled model collection comprises: (1) slow-thinking models, optimized for reasoning tasks that require high accuracy; (2) two series of adaptive-thinking models, which dynamically adjust reasoning strategies based on input tasks to maximize efficiency across diverse scenarios; and (3) distilled reward models, which enable further reinforcement learning of reasoning models using distilled knowledge. Comprehensive evaluations across multiple benchmarks demonstrate both high inference efficiency and strong reasoning performance for these models, as well as the practical utility of distilled reward models. We further show that these models support industry practitioners by providing scalable training and inference functionalities on the Alibaba Cloud PAI (Platform for Artificial Intelligence) platform.
title Thinking with DistilQwen: A Tale of Four Distilled Reasoning and Reward Model Series
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2511.01354