SCOPE: Signal-Calibrated On-Policy Distillation Enhancement with Dual-Path Adaptive Weighting

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zheng, Binbin, Ma, Xing, Liang, Yiheng, Ruan, Jingqing, Fu, Xiaoliang, Lin, Kepeng, Zhu, Benchang, Zeng, Ke, Cai, Xunliang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916070240026624
author Zheng, Binbin
Ma, Xing
Liang, Yiheng
Ruan, Jingqing
Fu, Xiaoliang
Lin, Kepeng
Zhu, Benchang
Zeng, Ke
Cai, Xunliang
author_facet Zheng, Binbin
Ma, Xing
Liang, Yiheng
Ruan, Jingqing
Fu, Xiaoliang
Lin, Kepeng
Zhu, Benchang
Zeng, Ke
Cai, Xunliang
contents On-policy reinforcement learning has become the dominant paradigm for reasoning alignment in large language models, yet its sparse, outcome-level rewards make token-level credit assignment notoriously difficult. On-Policy Distillation (OPD) alleviates this by introducing dense, token-level KL supervision from a teacher model, but typically applies this supervision uniformly across all rollouts, ignoring fundamental differences in signal quality. We propose Signal-Calibrated On-Policy Distillation Enhancement (SCOPE), a dual-path adaptive training framework that routes on-policy rollouts by correctness into two complementary supervision paths. For incorrect trajectories, SCOPE performs teacher-perplexity-weighted KL distillation to prioritize instances where the teacher demonstrates genuine corrective capability, while down-weighting unreliable guidance. For correct trajectories, it applies student-perplexity-weighted MLE to concentrate reinforcement on low-confidence samples at the capability boundary rather than over-reinforcing already mastered ones. Both paths employ a group-level normalization to adaptively calibrate weight distributions, accounting for the intrinsic difficulty variance across prompts. Extensive experiments on six reasoning benchmarks show that SCOPE achieves an average relative improvement of 11.42% in Avg@32 and 7.30% in Pass@32 over competitive baselines, demonstrating its consistent effectiveness.
format Preprint
id arxiv_https___arxiv_org_abs_2604_10688
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SCOPE: Signal-Calibrated On-Policy Distillation Enhancement with Dual-Path Adaptive Weighting
Zheng, Binbin
Ma, Xing
Liang, Yiheng
Ruan, Jingqing
Fu, Xiaoliang
Lin, Kepeng
Zhu, Benchang
Zeng, Ke
Cai, Xunliang
Machine Learning
Artificial Intelligence
Computation and Language
On-policy reinforcement learning has become the dominant paradigm for reasoning alignment in large language models, yet its sparse, outcome-level rewards make token-level credit assignment notoriously difficult. On-Policy Distillation (OPD) alleviates this by introducing dense, token-level KL supervision from a teacher model, but typically applies this supervision uniformly across all rollouts, ignoring fundamental differences in signal quality. We propose Signal-Calibrated On-Policy Distillation Enhancement (SCOPE), a dual-path adaptive training framework that routes on-policy rollouts by correctness into two complementary supervision paths. For incorrect trajectories, SCOPE performs teacher-perplexity-weighted KL distillation to prioritize instances where the teacher demonstrates genuine corrective capability, while down-weighting unreliable guidance. For correct trajectories, it applies student-perplexity-weighted MLE to concentrate reinforcement on low-confidence samples at the capability boundary rather than over-reinforcing already mastered ones. Both paths employ a group-level normalization to adaptively calibrate weight distributions, accounting for the intrinsic difficulty variance across prompts. Extensive experiments on six reasoning benchmarks show that SCOPE achieves an average relative improvement of 11.42% in Avg@32 and 7.30% in Pass@32 over competitive baselines, demonstrating its consistent effectiveness.
title SCOPE: Signal-Calibrated On-Policy Distillation Enhancement with Dual-Path Adaptive Weighting
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2604.10688