From Generic Correlation to Input-Specific Credit in On-Policy Self Distillation
Fuente:
arXiv
Salvato in:
| Autori principali: | Shen, Guobin, Huang, Lei, Cheng, Xiang, Zhao, Chenxiao, Li, Jindong, Zhao, Dongcheng, Yu, Xing |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Anti-Self-Distillation for Reasoning RL via Pointwise Mutual Information
di: Shen, Guobin, et al.
Pubblicazione: (2026)
di: Shen, Guobin, et al.
Pubblicazione: (2026)
VESPO: Variational Sequence-Level Soft Policy Optimization for Stable Off-Policy LLM Training
di: Shen, Guobin, et al.
Pubblicazione: (2026)
di: Shen, Guobin, et al.
Pubblicazione: (2026)
Multi-Level Safety Continual Projection for Fine-Tuned Large Language Models without Retraining
di: Han, Bing, et al.
Pubblicazione: (2025)
di: Han, Bing, et al.
Pubblicazione: (2025)
Light Alignment Improves LLM Safety via Model Self-Reflection with a Single Neuron
di: Shen, Sicheng, et al.
Pubblicazione: (2026)
di: Shen, Sicheng, et al.
Pubblicazione: (2026)
Safety Instincts: LLMs Learn to Trust Their Internal Compass for Self-Defense
di: Shen, Guobin, et al.
Pubblicazione: (2025)
di: Shen, Guobin, et al.
Pubblicazione: (2025)
OISD: On-Policy Internal Self-Distillation of Language Models
di: Liu, Xinyu, et al.
Pubblicazione: (2026)
di: Liu, Xinyu, et al.
Pubblicazione: (2026)
Self-ReSET: Learning to Self-Recover from Unsafe Reasoning Trajectories
di: Zhang, Dongcheng, et al.
Pubblicazione: (2026)
di: Zhang, Dongcheng, et al.
Pubblicazione: (2026)
PACED: Distillation and On-Policy Self-Distillation at the Frontier of Student Competence
di: Xu, Yuanda, et al.
Pubblicazione: (2026)
di: Xu, Yuanda, et al.
Pubblicazione: (2026)
Online Policy Distillation with Decision-Attention
di: Yu, Xinqiang, et al.
Pubblicazione: (2024)
di: Yu, Xinqiang, et al.
Pubblicazione: (2024)
HDPO: Hybrid Distillation Policy Optimization via Privileged Self-Distillation
di: Ding, Ken
Pubblicazione: (2026)
di: Ding, Ken
Pubblicazione: (2026)
A General Framework for Learning from Weak Supervision
di: Chen, Hao, et al.
Pubblicazione: (2024)
di: Chen, Hao, et al.
Pubblicazione: (2024)
OPD+: Rethinking the Advantage Design for On-Policy Distillation
di: Zhao, Hanyang, et al.
Pubblicazione: (2026)
di: Zhao, Hanyang, et al.
Pubblicazione: (2026)
Towards Reliable Evaluation of Adversarial Robustness for Spiking Neural Networks
di: Wang, Jihang, et al.
Pubblicazione: (2025)
di: Wang, Jihang, et al.
Pubblicazione: (2025)
On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes
di: Agarwal, Rishabh, et al.
Pubblicazione: (2023)
di: Agarwal, Rishabh, et al.
Pubblicazione: (2023)
Beyond Uniform Credit: Causal Credit Assignment for Policy Optimization
di: Khandoga, Mykola, et al.
Pubblicazione: (2026)
di: Khandoga, Mykola, et al.
Pubblicazione: (2026)
Implicit Federated In-context Learning For Task-Specific LLM Fine-Tuning
di: Li, Dongcheng, et al.
Pubblicazione: (2025)
di: Li, Dongcheng, et al.
Pubblicazione: (2025)
Graph Dimension Attention Networks for Enterprise Credit Assessment
di: Wei, Shaopeng, et al.
Pubblicazione: (2024)
di: Wei, Shaopeng, et al.
Pubblicazione: (2024)
Multimodal Generative Engine Optimization: Rank Manipulation for Vision-Language Model Rankers
di: Du, Yixuan, et al.
Pubblicazione: (2026)
di: Du, Yixuan, et al.
Pubblicazione: (2026)
Self-Supervised Quantization-Aware Knowledge Distillation
di: Zhao, Kaiqi, et al.
Pubblicazione: (2024)
di: Zhao, Kaiqi, et al.
Pubblicazione: (2024)
FedSDR: Federated Self-Distillation with Rectification
di: Ren, Ziheng, et al.
Pubblicazione: (2026)
di: Ren, Ziheng, et al.
Pubblicazione: (2026)
Bootstrapping Exploration with Group-Level Natural Language Feedback in Reinforcement Learning
di: Huang, Lei, et al.
Pubblicazione: (2026)
di: Huang, Lei, et al.
Pubblicazione: (2026)
Proximal Policy Distillation
di: Spigler, Giacomo
Pubblicazione: (2024)
di: Spigler, Giacomo
Pubblicazione: (2024)
When Are Teacher Tokens Reliable? Position-Weighted On-Policy Self-Distillation for Reasoning
di: Liu, Xiaogeng, et al.
Pubblicazione: (2026)
di: Liu, Xiaogeng, et al.
Pubblicazione: (2026)
Unifying Group-Relative and Self-Distillation Policy Optimization via Sample Routing
di: Li, Gengsheng, et al.
Pubblicazione: (2026)
di: Li, Gengsheng, et al.
Pubblicazione: (2026)
TRACE: Distilling Where It Matters via Token-Routed Self On-Policy Alignment
di: Wang, Jiaxuan, et al.
Pubblicazione: (2026)
di: Wang, Jiaxuan, et al.
Pubblicazione: (2026)
Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes
di: Fu, Yuqian, et al.
Pubblicazione: (2026)
di: Fu, Yuqian, et al.
Pubblicazione: (2026)
Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level
di: Jia, Nan, et al.
Pubblicazione: (2026)
di: Jia, Nan, et al.
Pubblicazione: (2026)
PromptBench: A Unified Library for Evaluation of Large Language Models
di: Zhu, Kaijie, et al.
Pubblicazione: (2023)
di: Zhu, Kaijie, et al.
Pubblicazione: (2023)
Dynamic Evaluation of Large Language Models by Meta Probing Agents
di: Zhu, Kaijie, et al.
Pubblicazione: (2024)
di: Zhu, Kaijie, et al.
Pubblicazione: (2024)
Prompt Tuning with Diffusion for Few-Shot Pre-trained Policy Generalization
di: Hu, Shengchao, et al.
Pubblicazione: (2024)
di: Hu, Shengchao, et al.
Pubblicazione: (2024)
DGPO: Distribution Guided Policy Optimization for Fine Grained Credit Assignment
di: Jin, Hongbo, et al.
Pubblicazione: (2026)
di: Jin, Hongbo, et al.
Pubblicazione: (2026)
Generalized Policy Gradient with History-Aware Decision Transformer for Reliable Routing over Graph Signals
di: Wei, Xing, et al.
Pubblicazione: (2025)
di: Wei, Xing, et al.
Pubblicazione: (2025)
BPL: Bias-adaptive Preference Distillation Learning for Recommender System
di: Kang, SeongKu, et al.
Pubblicazione: (2025)
di: Kang, SeongKu, et al.
Pubblicazione: (2025)
OGLS-SD: On-Policy Self-Distillation with Outcome-Guided Logit Steering for LLM Reasoning
di: Yang, Yuxiao, et al.
Pubblicazione: (2026)
di: Yang, Yuxiao, et al.
Pubblicazione: (2026)
Provable and Practical In-Context Policy Optimization for Self-Improvement
di: Yu, Tianrun, et al.
Pubblicazione: (2026)
di: Yu, Tianrun, et al.
Pubblicazione: (2026)
Annealing Self-Distillation Rectification Improves Adversarial Training
di: Wu, Yu-Yu, et al.
Pubblicazione: (2023)
di: Wu, Yu-Yu, et al.
Pubblicazione: (2023)
Model-Based Proactive Cost Generation for Learning Safe Policies Offline with Limited Violation Data
di: Xue, Ruiqi, et al.
Pubblicazione: (2026)
di: Xue, Ruiqi, et al.
Pubblicazione: (2026)
Extreme Region Policy Distillation
di: Chen, Changyu, et al.
Pubblicazione: (2026)
di: Chen, Changyu, et al.
Pubblicazione: (2026)
Towards Better Generalization via Distributional Input Projection Network
di: Hao, Yifan, et al.
Pubblicazione: (2025)
di: Hao, Yifan, et al.
Pubblicazione: (2025)
Self-Distillation for Multi-Token Prediction
di: Zhao, Guoliang, et al.
Pubblicazione: (2026)
di: Zhao, Guoliang, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Anti-Self-Distillation for Reasoning RL via Pointwise Mutual Information
di: Shen, Guobin, et al.
Pubblicazione: (2026) -
VESPO: Variational Sequence-Level Soft Policy Optimization for Stable Off-Policy LLM Training
di: Shen, Guobin, et al.
Pubblicazione: (2026) -
Multi-Level Safety Continual Projection for Fine-Tuned Large Language Models without Retraining
di: Han, Bing, et al.
Pubblicazione: (2025) -
Light Alignment Improves LLM Safety via Model Self-Reflection with a Single Neuron
di: Shen, Sicheng, et al.
Pubblicazione: (2026) -
Safety Instincts: LLMs Learn to Trust Their Internal Compass for Self-Defense
di: Shen, Guobin, et al.
Pubblicazione: (2025)