Scaling Reasoning Efficiently via Relaxed On-Policy Distillation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Ko, Jongwoo, Abdali, Sara, Kim, Young Jin, Chen, Tianyi, Cameron, Pashmina
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910049546272768
author Ko, Jongwoo
Abdali, Sara
Kim, Young Jin
Chen, Tianyi
Cameron, Pashmina
author_facet Ko, Jongwoo
Abdali, Sara
Kim, Young Jin
Chen, Tianyi
Cameron, Pashmina
contents On-policy distillation is pivotal for transferring reasoning capabilities to capacity-constrained models, yet remains prone to instability and negative transfer. We show that on-policy distillation can be interpreted, both theoretically and empirically, as a form of policy optimization, where the teacher-student log-likelihood ratio acts as a token reward. From this insight, we introduce REOPOLD (Relaxed On-Policy Distillation) a framework that stabilizes optimization by relaxing the strict imitation constraints of standard on-policy distillation. Specifically, REOPOLD temperately and selectively leverages rewards from the teacher through mixture-based reward clipping, entropy-based token-level dynamic sampling, and a unified exploration-to-refinement training strategy. Empirically, REOPOLD surpasses its baselines with superior sample efficiency during training and enhanced test-time scaling at inference, across mathematical, visual, and agentic tool-use reasoning tasks. Specifically, REOPOLD outperforms recent RL approaches achieving 6.7~12x greater sample efficiency and enables a 7B student to match a 32B teacher in visual reasoning with a ~3.32x inference speedup.
format Preprint
id arxiv_https___arxiv_org_abs_2603_11137
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Scaling Reasoning Efficiently via Relaxed On-Policy Distillation
Ko, Jongwoo
Abdali, Sara
Kim, Young Jin
Chen, Tianyi
Cameron, Pashmina
Machine Learning
Computation and Language
On-policy distillation is pivotal for transferring reasoning capabilities to capacity-constrained models, yet remains prone to instability and negative transfer. We show that on-policy distillation can be interpreted, both theoretically and empirically, as a form of policy optimization, where the teacher-student log-likelihood ratio acts as a token reward. From this insight, we introduce REOPOLD (Relaxed On-Policy Distillation) a framework that stabilizes optimization by relaxing the strict imitation constraints of standard on-policy distillation. Specifically, REOPOLD temperately and selectively leverages rewards from the teacher through mixture-based reward clipping, entropy-based token-level dynamic sampling, and a unified exploration-to-refinement training strategy. Empirically, REOPOLD surpasses its baselines with superior sample efficiency during training and enhanced test-time scaling at inference, across mathematical, visual, and agentic tool-use reasoning tasks. Specifically, REOPOLD outperforms recent RL approaches achieving 6.7~12x greater sample efficiency and enables a 7B student to match a 32B teacher in visual reasoning with a ~3.32x inference speedup.
title Scaling Reasoning Efficiently via Relaxed On-Policy Distillation
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2603.11137