When Are Teacher Tokens Reliable? Position-Weighted On-Policy Self-Distillation for Reasoning
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Xiaogeng, Wang, Xinyan, Ma, Yingzi, Zhang, Yechao, Xiao, Chaowei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ROM: Real-time Overthinking Mitigation via Streaming Detection and Intervention
di: Wang, Xinyan, et al.
Pubblicazione: (2026)
di: Wang, Xinyan, et al.
Pubblicazione: (2026)
ReasoningBomb: A Stealthy Denial-of-Service Attack by Inducing Pathologically Long Reasoning in Large Reasoning Models
di: Liu, Xiaogeng, et al.
Pubblicazione: (2026)
di: Liu, Xiaogeng, et al.
Pubblicazione: (2026)
AutoDAN-Reasoning: Enhancing Strategies Exploration based Jailbreak Attacks with Test-Time Scaling
di: Liu, Xiaogeng, et al.
Pubblicazione: (2025)
di: Liu, Xiaogeng, et al.
Pubblicazione: (2025)
Prune-OPD: Efficient and Reliable On-Policy Distillation for Long-Horizon Reasoning
di: Yang, Zhicheng, et al.
Pubblicazione: (2026)
di: Yang, Zhicheng, et al.
Pubblicazione: (2026)
TIP: Token Importance in On-Policy Distillation
di: Xu, Yuanda, et al.
Pubblicazione: (2026)
di: Xu, Yuanda, et al.
Pubblicazione: (2026)
TRACE: Distilling Where It Matters via Token-Routed Self On-Policy Alignment
di: Wang, Jiaxuan, et al.
Pubblicazione: (2026)
di: Wang, Jiaxuan, et al.
Pubblicazione: (2026)
Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level
di: Jia, Nan, et al.
Pubblicazione: (2026)
di: Jia, Nan, et al.
Pubblicazione: (2026)
Restoring the Sweet Spot: Pass-Rate Weighted Self-Distillation for LLM Reasoning
di: Liu, Zehao, et al.
Pubblicazione: (2026)
di: Liu, Zehao, et al.
Pubblicazione: (2026)
OGLS-SD: On-Policy Self-Distillation with Outcome-Guided Logit Steering for LLM Reasoning
di: Yang, Yuxiao, et al.
Pubblicazione: (2026)
di: Yang, Yuxiao, et al.
Pubblicazione: (2026)
MetaAgent: Automatically Constructing Multi-Agent Systems Based on Finite State Machines
di: Zhang, Yaolun, et al.
Pubblicazione: (2025)
di: Zhang, Yaolun, et al.
Pubblicazione: (2025)
PACED: Distillation and On-Policy Self-Distillation at the Frontier of Student Competence
di: Xu, Yuanda, et al.
Pubblicazione: (2026)
di: Xu, Yuanda, et al.
Pubblicazione: (2026)
AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs
di: Liu, Xiaogeng, et al.
Pubblicazione: (2024)
di: Liu, Xiaogeng, et al.
Pubblicazione: (2024)
AlignDistil: Token-Level Language Model Alignment as Adaptive Policy Distillation
di: Zhang, Songming, et al.
Pubblicazione: (2025)
di: Zhang, Songming, et al.
Pubblicazione: (2025)
Self-Distillation for Multi-Token Prediction
di: Zhao, Guoliang, et al.
Pubblicazione: (2026)
di: Zhao, Guoliang, et al.
Pubblicazione: (2026)
RePD: Defending Jailbreak Attack through a Retrieval-based Prompt Decomposition Process
di: Wang, Peiran, et al.
Pubblicazione: (2024)
di: Wang, Peiran, et al.
Pubblicazione: (2024)
HDPO: Hybrid Distillation Policy Optimization via Privileged Self-Distillation
di: Ding, Ken
Pubblicazione: (2026)
di: Ding, Ken
Pubblicazione: (2026)
Learning Diverse Policies with Soft Self-Generated Guidance
di: Wang, Guojian, et al.
Pubblicazione: (2024)
di: Wang, Guojian, et al.
Pubblicazione: (2024)
Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
di: Yang, Wenkai, et al.
Pubblicazione: (2026)
di: Yang, Wenkai, et al.
Pubblicazione: (2026)
One-Token Verification for Reasoning Correctness Estimation
di: Zhuang, Zhan, et al.
Pubblicazione: (2026)
di: Zhuang, Zhan, et al.
Pubblicazione: (2026)
Unifying Group-Relative and Self-Distillation Policy Optimization via Sample Routing
di: Li, Gengsheng, et al.
Pubblicazione: (2026)
di: Li, Gengsheng, et al.
Pubblicazione: (2026)
OET: Optimization-based prompt injection Evaluation Toolkit
di: Pan, Jinsheng, et al.
Pubblicazione: (2025)
di: Pan, Jinsheng, et al.
Pubblicazione: (2025)
SCOPE: Signal-Calibrated On-Policy Distillation Enhancement with Dual-Path Adaptive Weighting
di: Zheng, Binbin, et al.
Pubblicazione: (2026)
di: Zheng, Binbin, et al.
Pubblicazione: (2026)
Multi-Teacher Knowledge Distillation via Teacher-Informed Mixture Priors
di: Fang, Luyang, et al.
Pubblicazione: (2026)
di: Fang, Luyang, et al.
Pubblicazione: (2026)
AVSD: Adaptive-View Self-Distillation by Balancing Consensus and Teacher-Specific Privileged Signals
di: Nguyen, Duy, et al.
Pubblicazione: (2026)
di: Nguyen, Duy, et al.
Pubblicazione: (2026)
Extreme Region Policy Distillation
di: Chen, Changyu, et al.
Pubblicazione: (2026)
di: Chen, Changyu, et al.
Pubblicazione: (2026)
ERPO: Token-Level Entropy-Regulated Policy Optimization for Large Reasoning Models
di: Yu, Song, et al.
Pubblicazione: (2026)
di: Yu, Song, et al.
Pubblicazione: (2026)
wd1: Weighted Policy Optimization for Reasoning in Diffusion Language Models
di: Tang, Xiaohang, et al.
Pubblicazione: (2025)
di: Tang, Xiaohang, et al.
Pubblicazione: (2025)
Post-Training is About States, Not Tokens: A State Distribution View of SFT, RL, and On-Policy Distillation
di: Nie, Dong
Pubblicazione: (2026)
di: Nie, Dong
Pubblicazione: (2026)
Validity-Calibrated Reasoning Distillation
di: Saadi, Khouloud, et al.
Pubblicazione: (2026)
di: Saadi, Khouloud, et al.
Pubblicazione: (2026)
From Generic Correlation to Input-Specific Credit in On-Policy Self Distillation
di: Shen, Guobin, et al.
Pubblicazione: (2026)
di: Shen, Guobin, et al.
Pubblicazione: (2026)
ThinkSwitch: Context Distillation with LoRA and Weight Interpolation for Specific-Purpose Reasoning Tasks
di: Saini, Dhruv, et al.
Pubblicazione: (2026)
di: Saini, Dhruv, et al.
Pubblicazione: (2026)
HEAL: Hindsight Entropy-Assisted Learning for Reasoning Distillation
di: Zhang, Wenjing, et al.
Pubblicazione: (2026)
di: Zhang, Wenjing, et al.
Pubblicazione: (2026)
Toward Student-Oriented Teacher Network Training For Knowledge Distillation
di: Dong, Chengyu, et al.
Pubblicazione: (2022)
di: Dong, Chengyu, et al.
Pubblicazione: (2022)
CoVeR: Conformal Calibration for Versatile and Reliable Autoregressive Next-Token Prediction
di: Chen, Yuzhu, et al.
Pubblicazione: (2025)
di: Chen, Yuzhu, et al.
Pubblicazione: (2025)
BudgetThinker: Empowering Budget-aware LLM Reasoning with Control Tokens
di: Wen, Hao, et al.
Pubblicazione: (2025)
di: Wen, Hao, et al.
Pubblicazione: (2025)
TED: Training-Free Experience Distillation for Multimodal Reasoning
di: Yuan, Shuozhi, et al.
Pubblicazione: (2026)
di: Yuan, Shuozhi, et al.
Pubblicazione: (2026)
Proximal Policy Distillation
di: Spigler, Giacomo
Pubblicazione: (2024)
di: Spigler, Giacomo
Pubblicazione: (2024)
Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation
di: Wu, Yecheng, et al.
Pubblicazione: (2026)
di: Wu, Yecheng, et al.
Pubblicazione: (2026)
OISD: On-Policy Internal Self-Distillation of Language Models
di: Liu, Xinyu, et al.
Pubblicazione: (2026)
di: Liu, Xinyu, et al.
Pubblicazione: (2026)
Robust Knowledge Distillation Based on Feature Variance Against Backdoored Teacher Model
di: Chen, Jinyin, et al.
Pubblicazione: (2024)
di: Chen, Jinyin, et al.
Pubblicazione: (2024)
Documenti analoghi
-
ROM: Real-time Overthinking Mitigation via Streaming Detection and Intervention
di: Wang, Xinyan, et al.
Pubblicazione: (2026) -
ReasoningBomb: A Stealthy Denial-of-Service Attack by Inducing Pathologically Long Reasoning in Large Reasoning Models
di: Liu, Xiaogeng, et al.
Pubblicazione: (2026) -
AutoDAN-Reasoning: Enhancing Strategies Exploration based Jailbreak Attacks with Test-Time Scaling
di: Liu, Xiaogeng, et al.
Pubblicazione: (2025) -
Prune-OPD: Efficient and Reliable On-Policy Distillation for Long-Horizon Reasoning
di: Yang, Zhicheng, et al.
Pubblicazione: (2026) -
TIP: Token Importance in On-Policy Distillation
di: Xu, Yuanda, et al.
Pubblicazione: (2026)