Unveiling the Key Factors for Distilling Chain-of-Thought Reasoning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Chen, Xinghao, Sun, Zhijing, Guo, Wenjin, Zhang, Miaoran, Chen, Yanjun, Sun, Yirong, Su, Hui, Pan, Yijie, Klakow, Dietrich, Li, Wenjie, Shen, Xiaoyu
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913859857547264
author Chen, Xinghao
Sun, Zhijing
Guo, Wenjin
Zhang, Miaoran
Chen, Yanjun
Sun, Yirong
Su, Hui
Pan, Yijie
Klakow, Dietrich
Li, Wenjie
Shen, Xiaoyu
author_facet Chen, Xinghao
Sun, Zhijing
Guo, Wenjin
Zhang, Miaoran
Chen, Yanjun
Sun, Yirong
Su, Hui
Pan, Yijie
Klakow, Dietrich
Li, Wenjie
Shen, Xiaoyu
contents Large Language Models (LLMs) excel in reasoning tasks through Chain-of-Thought (CoT) prompting. However, CoT prompting greatly increases computational demands, which has prompted growing interest in distilling CoT capabilities into Small Language Models (SLMs). This study systematically examines the factors influencing CoT distillation, including the choice of granularity, format and teacher model. Through experiments involving four teacher models and seven student models across seven mathematical and commonsense reasoning datasets, we uncover three key findings: (1) Unlike LLMs, SLMs exhibit a non-monotonic relationship with granularity, with stronger models benefiting from finer-grained reasoning and weaker models performing better with simpler CoT supervision; (2) CoT format significantly impacts LLMs but has minimal effect on SLMs, likely due to their reliance on supervised fine-tuning rather than pretraining preferences; (3) Stronger teacher models do NOT always produce better student models, as diversity and complexity in CoT supervision can outweigh accuracy alone. These findings emphasize the need to tailor CoT strategies to specific student model, offering actionable insights for optimizing CoT distillation in SLMs. The code and datasets are available at https://github.com/EIT-NLP/Distilling-CoT-Reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2502_18001
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Unveiling the Key Factors for Distilling Chain-of-Thought Reasoning
Chen, Xinghao
Sun, Zhijing
Guo, Wenjin
Zhang, Miaoran
Chen, Yanjun
Sun, Yirong
Su, Hui
Pan, Yijie
Klakow, Dietrich
Li, Wenjie
Shen, Xiaoyu
Computation and Language
Large Language Models (LLMs) excel in reasoning tasks through Chain-of-Thought (CoT) prompting. However, CoT prompting greatly increases computational demands, which has prompted growing interest in distilling CoT capabilities into Small Language Models (SLMs). This study systematically examines the factors influencing CoT distillation, including the choice of granularity, format and teacher model. Through experiments involving four teacher models and seven student models across seven mathematical and commonsense reasoning datasets, we uncover three key findings: (1) Unlike LLMs, SLMs exhibit a non-monotonic relationship with granularity, with stronger models benefiting from finer-grained reasoning and weaker models performing better with simpler CoT supervision; (2) CoT format significantly impacts LLMs but has minimal effect on SLMs, likely due to their reliance on supervised fine-tuning rather than pretraining preferences; (3) Stronger teacher models do NOT always produce better student models, as diversity and complexity in CoT supervision can outweigh accuracy alone. These findings emphasize the need to tailor CoT strategies to specific student model, offering actionable insights for optimizing CoT distillation in SLMs. The code and datasets are available at https://github.com/EIT-NLP/Distilling-CoT-Reasoning.
title Unveiling the Key Factors for Distilling Chain-of-Thought Reasoning
topic Computation and Language
url https://arxiv.org/abs/2502.18001