The Signal is in the Steps: Local Scoring for Reasoning Data Selection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Just, Hoang Anh, Ko, Myeongseob, Jia, Ruoxi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910129697325056
author Just, Hoang Anh
Ko, Myeongseob
Jia, Ruoxi
author_facet Just, Hoang Anh
Ko, Myeongseob
Jia, Ruoxi
contents Distilling long-form reasoning from teacher models into smaller students requires selecting which candidate solutions to train on. Recent work argues that one should select responses the student model assigns highest probability, i.e., favoring solutions ``natural'' to the student. However, we find that this approach works within a single teacher but fails when scaling to long reasoning traces from multiple diverse teachers. We identify a key cause: this approach scores entire solutions, but students generalize by recombining familiar reasoning steps, not by memorizing complete solutions. Full-trajectory scoring optimizes the wrong target; it rewards global fluency while the transferable signal lies in local step transitions. We propose Local Average Log Probability (LALP), which scores each reasoning step using only a small window of preceding context, measuring whether each step is justified by its immediate premises rather than whether the full response looks natural to the student. LALP enables two practical use cases: selecting the best teacher before fine-tuning and curating training data from diverse teacher pools. Across math, coding, and science reasoning tasks, LALP consistently improves accuracy when selecting the most natural solutions by a large margin.
format Preprint
id arxiv_https___arxiv_org_abs_2510_03988
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Signal is in the Steps: Local Scoring for Reasoning Data Selection
Just, Hoang Anh
Ko, Myeongseob
Jia, Ruoxi
Machine Learning
Artificial Intelligence
Distilling long-form reasoning from teacher models into smaller students requires selecting which candidate solutions to train on. Recent work argues that one should select responses the student model assigns highest probability, i.e., favoring solutions ``natural'' to the student. However, we find that this approach works within a single teacher but fails when scaling to long reasoning traces from multiple diverse teachers. We identify a key cause: this approach scores entire solutions, but students generalize by recombining familiar reasoning steps, not by memorizing complete solutions. Full-trajectory scoring optimizes the wrong target; it rewards global fluency while the transferable signal lies in local step transitions. We propose Local Average Log Probability (LALP), which scores each reasoning step using only a small window of preceding context, measuring whether each step is justified by its immediate premises rather than whether the full response looks natural to the student. LALP enables two practical use cases: selecting the best teacher before fine-tuning and curating training data from diverse teacher pools. Across math, coding, and science reasoning tasks, LALP consistently improves accuracy when selecting the most natural solutions by a large margin.
title The Signal is in the Steps: Local Scoring for Reasoning Data Selection
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2510.03988