Surgical Post-Training: Proximal On-Policy Distillation for Reasoning with Knowledge Retention
Fuente:
arXiv
Saved in:
| Main Authors: | Lin, Wenye, Han, Kai |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning
by: Xu, Hongling, et al.
Published: (2025)
by: Xu, Hongling, et al.
Published: (2025)
Updating Parametric Knowledge with Context Distillation Retains Post-Training Capabilities
by: Padmanabhan, Shankar, et al.
Published: (2026)
by: Padmanabhan, Shankar, et al.
Published: (2026)
Retentive or Forgetful? Diving into the Knowledge Memorizing Mechanism of Language Models
by: Cao, Boxi, et al.
Published: (2023)
by: Cao, Boxi, et al.
Published: (2023)
Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation
by: Yang, Zhuolin, et al.
Published: (2026)
by: Yang, Zhuolin, et al.
Published: (2026)
Dual-Head Reasoning Distillation: Improving Classifier Accuracy with Train-Time-Only Reasoning
by: Xu, Jillian, et al.
Published: (2025)
by: Xu, Jillian, et al.
Published: (2025)
Beyond Scaling: Measuring and Predicting the Upper Bound of Knowledge Retention in Language Model Pre-Training
by: Jiang, Changhao, et al.
Published: (2025)
by: Jiang, Changhao, et al.
Published: (2025)
Reliable Reasoning Path: Distilling Effective Guidance for LLM Reasoning with Knowledge Graphs
by: Xiao, Yilin, et al.
Published: (2025)
by: Xiao, Yilin, et al.
Published: (2025)
Bridging Reasoning Trajectories in On-Policy Distillation via Near-Future Guidance
by: Jiang, Yuxuan, et al.
Published: (2026)
by: Jiang, Yuxuan, et al.
Published: (2026)
Retention Consequence in Lifecycle Memory Control
by: Han, Jiarui
Published: (2026)
by: Han, Jiarui
Published: (2026)
A Primer in Post-Training Reasoning Data: What We Know About How It Works
by: Li, Yaoming, et al.
Published: (2026)
by: Li, Yaoming, et al.
Published: (2026)
Multimodal Commonsense Knowledge Distillation for Visual Question Answering
by: Yang, Shuo, et al.
Published: (2024)
by: Yang, Shuo, et al.
Published: (2024)
Efficient Post-Training Refinement of Latent Reasoning in Large Language Models
by: Wang, Xinyuan, et al.
Published: (2025)
by: Wang, Xinyuan, et al.
Published: (2025)
Thinking About Thinking: Evaluating Reasoning in Post-Trained Language Models
by: Singla, Pratham, et al.
Published: (2025)
by: Singla, Pratham, et al.
Published: (2025)
Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
by: Yang, Wenkai, et al.
Published: (2026)
by: Yang, Wenkai, et al.
Published: (2026)
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training
by: Pan, Chengjun, et al.
Published: (2026)
by: Pan, Chengjun, et al.
Published: (2026)
Learning to Reason Efficiently with A* Post-Training
by: Opedal, Andreas, et al.
Published: (2026)
by: Opedal, Andreas, et al.
Published: (2026)
Hybrid Policy Distillation for LLMs
by: Zhu, Wenhong, et al.
Published: (2026)
by: Zhu, Wenhong, et al.
Published: (2026)
Beyond Answers: Transferring Reasoning Capabilities to Smaller LLMs Using Multi-Teacher Knowledge Distillation
by: Tian, Yijun, et al.
Published: (2024)
by: Tian, Yijun, et al.
Published: (2024)
Distilling Reasoning Without Knowledge: A Framework for Reliable LLMs
by: Kietkajornrit, Auksarapak, et al.
Published: (2026)
by: Kietkajornrit, Auksarapak, et al.
Published: (2026)
Towards Understanding and Improving Knowledge Distillation for Neural Machine Translation
by: Zhang, Songming, et al.
Published: (2023)
by: Zhang, Songming, et al.
Published: (2023)
Beyond Perplexity: A Lightweight Benchmark for Knowledge Retention in Supervised Fine-Tuning
by: Shabgahi, Soheil Zibakhsh, et al.
Published: (2026)
by: Shabgahi, Soheil Zibakhsh, et al.
Published: (2026)
Towards Robust Pruning: An Adaptive Knowledge-Retention Pruning Strategy for Language Models
by: Li, Jianwei, et al.
Published: (2023)
by: Li, Jianwei, et al.
Published: (2023)
Can Post-Training Transform LLMs into Causal Reasoners?
by: Chen, Junqi, et al.
Published: (2026)
by: Chen, Junqi, et al.
Published: (2026)
Distilling the Essence: Efficient Reasoning Distillation via Sequence Truncation
by: Chen, Wei-Rui, et al.
Published: (2025)
by: Chen, Wei-Rui, et al.
Published: (2025)
Your Teacher Can't Help You Here: Combating Supervision Fidelity Decay in On-Policy Distillation
by: Liu, Yanjiang, et al.
Published: (2026)
by: Liu, Yanjiang, et al.
Published: (2026)
Memorizing is Not Enough: Deep Knowledge Injection Through Reasoning
by: Xu, Ruoxi, et al.
Published: (2025)
by: Xu, Ruoxi, et al.
Published: (2025)
SD-Search: On-Policy Hindsight Self-Distillation for Search-Augmented Reasoning
by: Ma, Yufei, et al.
Published: (2026)
by: Ma, Yufei, et al.
Published: (2026)
AdaSwitch: Balancing Exploration and Guidance in Knowledge Distillation via Adaptive Switching
by: Peng, Jingyu, et al.
Published: (2025)
by: Peng, Jingyu, et al.
Published: (2025)
From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning
by: Wang, Shaojie, et al.
Published: (2026)
by: Wang, Shaojie, et al.
Published: (2026)
Decoder-based Sense Knowledge Distillation
by: Wang, Qitong, et al.
Published: (2026)
by: Wang, Qitong, et al.
Published: (2026)
Self-Knowledge Distillation for Learning Ambiguity
by: Park, Hancheol, et al.
Published: (2024)
by: Park, Hancheol, et al.
Published: (2024)
Efficient Mathematical Reasoning Models via Dynamic Pruning and Knowledge Distillation
by: Yu, Fengming, et al.
Published: (2025)
by: Yu, Fengming, et al.
Published: (2025)
Training the Knowledge Base through Evidence Distillation and Write-Back Enrichment
by: Lu, Yuxing, et al.
Published: (2026)
by: Lu, Yuxing, et al.
Published: (2026)
STEPER: Step-wise Knowledge Distillation for Enhancing Reasoning Ability in Multi-Step Retrieval-Augmented Language Models
by: Lee, Kyumin, et al.
Published: (2025)
by: Lee, Kyumin, et al.
Published: (2025)
Efficient Technical Term Translation: A Knowledge Distillation Approach for Parenthetical Terminology Translation
by: Myung, Jiyoon, et al.
Published: (2024)
by: Myung, Jiyoon, et al.
Published: (2024)
Post-Trained MoE Can Skip Half Experts via Self-Distillation
by: Lv, Xingtai, et al.
Published: (2026)
by: Lv, Xingtai, et al.
Published: (2026)
ARise: Towards Knowledge-Augmented Reasoning via Risk-Adaptive Search
by: Zhang, Yize, et al.
Published: (2025)
by: Zhang, Yize, et al.
Published: (2025)
Large Language Model Post-Training: A Unified View of Off-Policy and On-Policy Learning
by: Zhao, Shiwan, et al.
Published: (2026)
by: Zhao, Shiwan, et al.
Published: (2026)
Knowledge Reasoning Language Model: Unifying Knowledge and Language for Inductive Knowledge Graph Reasoning
by: Zhuo, Xingrui, et al.
Published: (2025)
by: Zhuo, Xingrui, et al.
Published: (2025)
LLM-Oriented Token-Adaptive Knowledge Distillation
by: Xie, Xurong, et al.
Published: (2025)
by: Xie, Xurong, et al.
Published: (2025)
Similar Items
-
KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning
by: Xu, Hongling, et al.
Published: (2025) -
Updating Parametric Knowledge with Context Distillation Retains Post-Training Capabilities
by: Padmanabhan, Shankar, et al.
Published: (2026) -
Retentive or Forgetful? Diving into the Knowledge Memorizing Mechanism of Language Models
by: Cao, Boxi, et al.
Published: (2023) -
Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation
by: Yang, Zhuolin, et al.
Published: (2026) -
Dual-Head Reasoning Distillation: Improving Classifier Accuracy with Train-Time-Only Reasoning
by: Xu, Jillian, et al.
Published: (2025)