Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Jeonghye, Jeon, Jiwon, Li, Dongsheng, Yang, Yuqing
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909033330376704
author Kim, Jeonghye
Jeon, Jiwon
Li, Dongsheng
Yang, Yuqing
author_facet Kim, Jeonghye
Jeon, Jiwon
Li, Dongsheng
Yang, Yuqing
contents Self-distillation has emerged as a powerful framework for post-training LLMs, where a teacher conditioned on extra information guides a student without it, both from the same model. While this guidance is useful when the student has failed, on successful rollouts, the same mechanism instead overwrites the student's choices and suppresses it's own reasoning. Therefore, we propose reading the original self-distillation signal in reverse: when the student succeeds along a path the teacher would not have predicted, these tokens reflect its self-driven reasoning. Building on this, we propose RLRT (RLVR with Reversed Teacher), which augments GRPO by reinforcing these tokens on correct rollouts. We interpret this as a new form of exploration in RLVR: not uniform diversity, but valuable exploration grounded in the student's own success. Across base, instruction-tuned, and thinking-tuned Qwen3 checkpoints, RLRT substantially outperforms self-distillation and exploration-based baselines, establishing information asymmetry as a new, principled design axis for RLVR.
format Preprint
id arxiv_https___arxiv_org_abs_2605_10781
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR
Kim, Jeonghye
Jeon, Jiwon
Li, Dongsheng
Yang, Yuqing
Machine Learning
Computation and Language
Self-distillation has emerged as a powerful framework for post-training LLMs, where a teacher conditioned on extra information guides a student without it, both from the same model. While this guidance is useful when the student has failed, on successful rollouts, the same mechanism instead overwrites the student's choices and suppresses it's own reasoning. Therefore, we propose reading the original self-distillation signal in reverse: when the student succeeds along a path the teacher would not have predicted, these tokens reflect its self-driven reasoning. Building on this, we propose RLRT (RLVR with Reversed Teacher), which augments GRPO by reinforcing these tokens on correct rollouts. We interpret this as a new form of exploration in RLVR: not uniform diversity, but valuable exploration grounded in the student's own success. Across base, instruction-tuned, and thinking-tuned Qwen3 checkpoints, RLRT substantially outperforms self-distillation and exploration-based baselines, establishing information asymmetry as a new, principled design axis for RLVR.
title Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2605.10781