The Quest for Efficient Reasoning: A Data-Centric Benchmark to CoT Distillation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866915778703392768 |
|---|---|
| author | Zhang, Ruichen Khan, Rana Muhammad Shahroz Tan, Zhen Li, Dawei Wang, Song Chen, Tianlong |
| author_facet | Zhang, Ruichen Khan, Rana Muhammad Shahroz Tan, Zhen Li, Dawei Wang, Song Chen, Tianlong |
| contents | Data-centric distillation, including data augmentation, selection, and mixing, offers a promising path to creating smaller, more efficient student Large Language Models (LLMs) that retain strong reasoning abilities. However, there still lacks a comprehensive benchmark to systematically assess the effect of each distillation approach. This paper introduces DC-CoT, the first data-centric benchmark that investigates data manipulation in chain-of-thought (CoT) distillation from method, model and data perspectives. Utilizing various teacher models (e.g., o4-mini, Gemini-Pro, Claude-3.5) and student architectures (e.g., 3B, 7B parameters), we rigorously evaluate the impact of these data manipulations on student model performance across multiple reasoning datasets, with a focus on in-distribution (IID) and out-of-distribution (OOD) generalization, and cross-domain transfer. Our findings aim to provide actionable insights and establish best practices for optimizing CoT distillation through data-centric techniques, ultimately facilitating the development of more accessible and capable reasoning models. The codebase can be accessed at https://github.com/UNITES-Lab/Distillation-Bench |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_18759 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | The Quest for Efficient Reasoning: A Data-Centric Benchmark to CoT Distillation Zhang, Ruichen Khan, Rana Muhammad Shahroz Tan, Zhen Li, Dawei Wang, Song Chen, Tianlong Artificial Intelligence Machine Learning Data-centric distillation, including data augmentation, selection, and mixing, offers a promising path to creating smaller, more efficient student Large Language Models (LLMs) that retain strong reasoning abilities. However, there still lacks a comprehensive benchmark to systematically assess the effect of each distillation approach. This paper introduces DC-CoT, the first data-centric benchmark that investigates data manipulation in chain-of-thought (CoT) distillation from method, model and data perspectives. Utilizing various teacher models (e.g., o4-mini, Gemini-Pro, Claude-3.5) and student architectures (e.g., 3B, 7B parameters), we rigorously evaluate the impact of these data manipulations on student model performance across multiple reasoning datasets, with a focus on in-distribution (IID) and out-of-distribution (OOD) generalization, and cross-domain transfer. Our findings aim to provide actionable insights and establish best practices for optimizing CoT distillation through data-centric techniques, ultimately facilitating the development of more accessible and capable reasoning models. The codebase can be accessed at https://github.com/UNITES-Lab/Distillation-Bench |
| title | The Quest for Efficient Reasoning: A Data-Centric Benchmark to CoT Distillation |
| topic | Artificial Intelligence Machine Learning |
| url | https://arxiv.org/abs/2505.18759 |