Climbing the Ladder of Reasoning: What LLMs Can-and Still Can't-Solve after SFT?

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Sun, Yiyou, Zhou, Georgia, Bai, Haoyue, Wang, Hao, Li, Dacheng, Dziri, Nouha, Song, Dawn
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911363835625472
author Sun, Yiyou
Zhou, Georgia
Bai, Haoyue
Wang, Hao
Li, Dacheng
Dziri, Nouha
Song, Dawn
author_facet Sun, Yiyou
Zhou, Georgia
Bai, Haoyue
Wang, Hao
Li, Dacheng
Dziri, Nouha
Song, Dawn
contents Recent supervised fine-tuning (SFT) approaches have significantly improved language models' performance on mathematical reasoning tasks, even when models are trained at a small scale. However, the specific capabilities enhanced through such fine-tuning remain poorly understood. In this paper, we conduct a detailed analysis of model performance on the AIME24 dataset to understand how reasoning capabilities evolve. We discover a ladder-like structure in problem difficulty, categorize questions into four tiers (Easy, Medium, Hard, and Extremely Hard (Exh)), and identify the specific requirements for advancing between tiers. We find that progression from Easy to Medium tier requires adopting an R1 reasoning style with minimal SFT (500-1K instances), while Hard-level questions suffer from frequent model's errors at each step of the reasoning chain, with accuracy plateauing at around 65% despite logarithmic scaling. Exh-level questions present a fundamentally different challenge; they require unconventional problem-solving skills that current models uniformly struggle with. Additional findings reveal that carefully curated small-scale datasets offer limited advantage-scaling dataset size proves far more effective. Our analysis provides a clearer roadmap for advancing language model capabilities in mathematical reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2504_11741
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Climbing the Ladder of Reasoning: What LLMs Can-and Still Can't-Solve after SFT?
Sun, Yiyou
Zhou, Georgia
Bai, Haoyue
Wang, Hao
Li, Dacheng
Dziri, Nouha
Song, Dawn
Artificial Intelligence
Computation and Language
Machine Learning
Recent supervised fine-tuning (SFT) approaches have significantly improved language models' performance on mathematical reasoning tasks, even when models are trained at a small scale. However, the specific capabilities enhanced through such fine-tuning remain poorly understood. In this paper, we conduct a detailed analysis of model performance on the AIME24 dataset to understand how reasoning capabilities evolve. We discover a ladder-like structure in problem difficulty, categorize questions into four tiers (Easy, Medium, Hard, and Extremely Hard (Exh)), and identify the specific requirements for advancing between tiers. We find that progression from Easy to Medium tier requires adopting an R1 reasoning style with minimal SFT (500-1K instances), while Hard-level questions suffer from frequent model's errors at each step of the reasoning chain, with accuracy plateauing at around 65% despite logarithmic scaling. Exh-level questions present a fundamentally different challenge; they require unconventional problem-solving skills that current models uniformly struggle with. Additional findings reveal that carefully curated small-scale datasets offer limited advantage-scaling dataset size proves far more effective. Our analysis provides a clearer roadmap for advancing language model capabilities in mathematical reasoning.
title Climbing the Ladder of Reasoning: What LLMs Can-and Still Can't-Solve after SFT?
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2504.11741