From Generic Correlation to Input-Specific Credit in On-Policy Self Distillation
Fuente:
arXiv
Guardado en:
| Autores principales: | Shen, Guobin, Huang, Lei, Cheng, Xiang, Zhao, Chenxiao, Li, Jindong, Zhao, Dongcheng, Yu, Xing |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Anti-Self-Distillation for Reasoning RL via Pointwise Mutual Information
por: Shen, Guobin, et al.
Publicado: (2026)
por: Shen, Guobin, et al.
Publicado: (2026)
VESPO: Variational Sequence-Level Soft Policy Optimization for Stable Off-Policy LLM Training
por: Shen, Guobin, et al.
Publicado: (2026)
por: Shen, Guobin, et al.
Publicado: (2026)
Multi-Level Safety Continual Projection for Fine-Tuned Large Language Models without Retraining
por: Han, Bing, et al.
Publicado: (2025)
por: Han, Bing, et al.
Publicado: (2025)
Light Alignment Improves LLM Safety via Model Self-Reflection with a Single Neuron
por: Shen, Sicheng, et al.
Publicado: (2026)
por: Shen, Sicheng, et al.
Publicado: (2026)
Safety Instincts: LLMs Learn to Trust Their Internal Compass for Self-Defense
por: Shen, Guobin, et al.
Publicado: (2025)
por: Shen, Guobin, et al.
Publicado: (2025)
OISD: On-Policy Internal Self-Distillation of Language Models
por: Liu, Xinyu, et al.
Publicado: (2026)
por: Liu, Xinyu, et al.
Publicado: (2026)
Self-ReSET: Learning to Self-Recover from Unsafe Reasoning Trajectories
por: Zhang, Dongcheng, et al.
Publicado: (2026)
por: Zhang, Dongcheng, et al.
Publicado: (2026)
PACED: Distillation and On-Policy Self-Distillation at the Frontier of Student Competence
por: Xu, Yuanda, et al.
Publicado: (2026)
por: Xu, Yuanda, et al.
Publicado: (2026)
Online Policy Distillation with Decision-Attention
por: Yu, Xinqiang, et al.
Publicado: (2024)
por: Yu, Xinqiang, et al.
Publicado: (2024)
HDPO: Hybrid Distillation Policy Optimization via Privileged Self-Distillation
por: Ding, Ken
Publicado: (2026)
por: Ding, Ken
Publicado: (2026)
A General Framework for Learning from Weak Supervision
por: Chen, Hao, et al.
Publicado: (2024)
por: Chen, Hao, et al.
Publicado: (2024)
OPD+: Rethinking the Advantage Design for On-Policy Distillation
por: Zhao, Hanyang, et al.
Publicado: (2026)
por: Zhao, Hanyang, et al.
Publicado: (2026)
Towards Reliable Evaluation of Adversarial Robustness for Spiking Neural Networks
por: Wang, Jihang, et al.
Publicado: (2025)
por: Wang, Jihang, et al.
Publicado: (2025)
On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes
por: Agarwal, Rishabh, et al.
Publicado: (2023)
por: Agarwal, Rishabh, et al.
Publicado: (2023)
Beyond Uniform Credit: Causal Credit Assignment for Policy Optimization
por: Khandoga, Mykola, et al.
Publicado: (2026)
por: Khandoga, Mykola, et al.
Publicado: (2026)
Implicit Federated In-context Learning For Task-Specific LLM Fine-Tuning
por: Li, Dongcheng, et al.
Publicado: (2025)
por: Li, Dongcheng, et al.
Publicado: (2025)
Graph Dimension Attention Networks for Enterprise Credit Assessment
por: Wei, Shaopeng, et al.
Publicado: (2024)
por: Wei, Shaopeng, et al.
Publicado: (2024)
Multimodal Generative Engine Optimization: Rank Manipulation for Vision-Language Model Rankers
por: Du, Yixuan, et al.
Publicado: (2026)
por: Du, Yixuan, et al.
Publicado: (2026)
Self-Supervised Quantization-Aware Knowledge Distillation
por: Zhao, Kaiqi, et al.
Publicado: (2024)
por: Zhao, Kaiqi, et al.
Publicado: (2024)
FedSDR: Federated Self-Distillation with Rectification
por: Ren, Ziheng, et al.
Publicado: (2026)
por: Ren, Ziheng, et al.
Publicado: (2026)
Bootstrapping Exploration with Group-Level Natural Language Feedback in Reinforcement Learning
por: Huang, Lei, et al.
Publicado: (2026)
por: Huang, Lei, et al.
Publicado: (2026)
Proximal Policy Distillation
por: Spigler, Giacomo
Publicado: (2024)
por: Spigler, Giacomo
Publicado: (2024)
When Are Teacher Tokens Reliable? Position-Weighted On-Policy Self-Distillation for Reasoning
por: Liu, Xiaogeng, et al.
Publicado: (2026)
por: Liu, Xiaogeng, et al.
Publicado: (2026)
Unifying Group-Relative and Self-Distillation Policy Optimization via Sample Routing
por: Li, Gengsheng, et al.
Publicado: (2026)
por: Li, Gengsheng, et al.
Publicado: (2026)
TRACE: Distilling Where It Matters via Token-Routed Self On-Policy Alignment
por: Wang, Jiaxuan, et al.
Publicado: (2026)
por: Wang, Jiaxuan, et al.
Publicado: (2026)
Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes
por: Fu, Yuqian, et al.
Publicado: (2026)
por: Fu, Yuqian, et al.
Publicado: (2026)
Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level
por: Jia, Nan, et al.
Publicado: (2026)
por: Jia, Nan, et al.
Publicado: (2026)
PromptBench: A Unified Library for Evaluation of Large Language Models
por: Zhu, Kaijie, et al.
Publicado: (2023)
por: Zhu, Kaijie, et al.
Publicado: (2023)
Dynamic Evaluation of Large Language Models by Meta Probing Agents
por: Zhu, Kaijie, et al.
Publicado: (2024)
por: Zhu, Kaijie, et al.
Publicado: (2024)
Prompt Tuning with Diffusion for Few-Shot Pre-trained Policy Generalization
por: Hu, Shengchao, et al.
Publicado: (2024)
por: Hu, Shengchao, et al.
Publicado: (2024)
DGPO: Distribution Guided Policy Optimization for Fine Grained Credit Assignment
por: Jin, Hongbo, et al.
Publicado: (2026)
por: Jin, Hongbo, et al.
Publicado: (2026)
Generalized Policy Gradient with History-Aware Decision Transformer for Reliable Routing over Graph Signals
por: Wei, Xing, et al.
Publicado: (2025)
por: Wei, Xing, et al.
Publicado: (2025)
BPL: Bias-adaptive Preference Distillation Learning for Recommender System
por: Kang, SeongKu, et al.
Publicado: (2025)
por: Kang, SeongKu, et al.
Publicado: (2025)
OGLS-SD: On-Policy Self-Distillation with Outcome-Guided Logit Steering for LLM Reasoning
por: Yang, Yuxiao, et al.
Publicado: (2026)
por: Yang, Yuxiao, et al.
Publicado: (2026)
Provable and Practical In-Context Policy Optimization for Self-Improvement
por: Yu, Tianrun, et al.
Publicado: (2026)
por: Yu, Tianrun, et al.
Publicado: (2026)
Annealing Self-Distillation Rectification Improves Adversarial Training
por: Wu, Yu-Yu, et al.
Publicado: (2023)
por: Wu, Yu-Yu, et al.
Publicado: (2023)
Model-Based Proactive Cost Generation for Learning Safe Policies Offline with Limited Violation Data
por: Xue, Ruiqi, et al.
Publicado: (2026)
por: Xue, Ruiqi, et al.
Publicado: (2026)
Extreme Region Policy Distillation
por: Chen, Changyu, et al.
Publicado: (2026)
por: Chen, Changyu, et al.
Publicado: (2026)
Towards Better Generalization via Distributional Input Projection Network
por: Hao, Yifan, et al.
Publicado: (2025)
por: Hao, Yifan, et al.
Publicado: (2025)
Self-Distillation for Multi-Token Prediction
por: Zhao, Guoliang, et al.
Publicado: (2026)
por: Zhao, Guoliang, et al.
Publicado: (2026)
Ejemplares similares
-
Anti-Self-Distillation for Reasoning RL via Pointwise Mutual Information
por: Shen, Guobin, et al.
Publicado: (2026) -
VESPO: Variational Sequence-Level Soft Policy Optimization for Stable Off-Policy LLM Training
por: Shen, Guobin, et al.
Publicado: (2026) -
Multi-Level Safety Continual Projection for Fine-Tuned Large Language Models without Retraining
por: Han, Bing, et al.
Publicado: (2025) -
Light Alignment Improves LLM Safety via Model Self-Reflection with a Single Neuron
por: Shen, Sicheng, et al.
Publicado: (2026) -
Safety Instincts: LLMs Learn to Trust Their Internal Compass for Self-Defense
por: Shen, Guobin, et al.
Publicado: (2025)