LLMs Can Easily Learn to Reason from Demonstrations Structure, not content, is what matters!

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Li, Dacheng, Cao, Shiyi, Griggs, Tyler, Liu, Shu, Mo, Xiangxi, Tang, Eric, Hegde, Sumanth, Hakhamaneshi, Kourosh, Patil, Shishir G., Zaharia, Matei, Gonzalez, Joseph E., Stoica, Ion
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908230538493952
author Li, Dacheng
Cao, Shiyi
Griggs, Tyler
Liu, Shu
Mo, Xiangxi
Tang, Eric
Hegde, Sumanth
Hakhamaneshi, Kourosh
Patil, Shishir G.
Zaharia, Matei
Gonzalez, Joseph E.
Stoica, Ion
author_facet Li, Dacheng
Cao, Shiyi
Griggs, Tyler
Liu, Shu
Mo, Xiangxi
Tang, Eric
Hegde, Sumanth
Hakhamaneshi, Kourosh
Patil, Shishir G.
Zaharia, Matei
Gonzalez, Joseph E.
Stoica, Ion
contents Large reasoning models (LRMs) tackle complex reasoning problems by following long chain-of-thoughts (Long CoT) that incorporate reflection, backtracking, and self-validation. However, the training techniques and data requirements to elicit Long CoT remain poorly understood. In this work, we find that a Large Language model (LLM) can effectively learn Long CoT reasoning through data-efficient supervised fine-tuning (SFT) and parameter-efficient low-rank adaptation (LoRA). With just 17k long CoT training samples, the Qwen2.5-32B-Instruct model achieves significant improvements on a wide range of math and coding benchmarks, including 56.7% (+40.0%) on AIME 2024 and 57.0% (+8.1%) on LiveCodeBench, competitive to the proprietary o1-preview model's score of 44.6% and 59.1%. More importantly, we find that the structure of Long CoT is critical to the learning process, whereas the content of individual reasoning steps has minimal impact. Perturbations affecting content, such as training on incorrect samples or removing reasoning keywords, have little impact on performance. In contrast, structural modifications that disrupt logical consistency in the Long CoT, such as shuffling or deleting reasoning steps, significantly degrade accuracy. For example, a model trained on Long CoT samples with incorrect answers still achieves only 3.2% lower accuracy compared to training with fully correct samples. These insights deepen our understanding of how to elicit reasoning capabilities in LLMs and highlight key considerations for efficiently training the next generation of reasoning models. This is the academic paper of our previous released Sky-T1-32B-Preview model. Codes are available at https://github.com/NovaSky-AI/SkyThought.
format Preprint
id arxiv_https___arxiv_org_abs_2502_07374
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LLMs Can Easily Learn to Reason from Demonstrations Structure, not content, is what matters!
Li, Dacheng
Cao, Shiyi
Griggs, Tyler
Liu, Shu
Mo, Xiangxi
Tang, Eric
Hegde, Sumanth
Hakhamaneshi, Kourosh
Patil, Shishir G.
Zaharia, Matei
Gonzalez, Joseph E.
Stoica, Ion
Artificial Intelligence
Large reasoning models (LRMs) tackle complex reasoning problems by following long chain-of-thoughts (Long CoT) that incorporate reflection, backtracking, and self-validation. However, the training techniques and data requirements to elicit Long CoT remain poorly understood. In this work, we find that a Large Language model (LLM) can effectively learn Long CoT reasoning through data-efficient supervised fine-tuning (SFT) and parameter-efficient low-rank adaptation (LoRA). With just 17k long CoT training samples, the Qwen2.5-32B-Instruct model achieves significant improvements on a wide range of math and coding benchmarks, including 56.7% (+40.0%) on AIME 2024 and 57.0% (+8.1%) on LiveCodeBench, competitive to the proprietary o1-preview model's score of 44.6% and 59.1%. More importantly, we find that the structure of Long CoT is critical to the learning process, whereas the content of individual reasoning steps has minimal impact. Perturbations affecting content, such as training on incorrect samples or removing reasoning keywords, have little impact on performance. In contrast, structural modifications that disrupt logical consistency in the Long CoT, such as shuffling or deleting reasoning steps, significantly degrade accuracy. For example, a model trained on Long CoT samples with incorrect answers still achieves only 3.2% lower accuracy compared to training with fully correct samples. These insights deepen our understanding of how to elicit reasoning capabilities in LLMs and highlight key considerations for efficiently training the next generation of reasoning models. This is the academic paper of our previous released Sky-T1-32B-Preview model. Codes are available at https://github.com/NovaSky-AI/SkyThought.
title LLMs Can Easily Learn to Reason from Demonstrations Structure, not content, is what matters!
topic Artificial Intelligence
url https://arxiv.org/abs/2502.07374