HDPO: Hybrid Distillation Policy Optimization via Privileged Self-Distillation
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Ding, Ken |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Privileged Information Distillation for Language Models
von: Penaloza, Emiliano, et al.
Veröffentlicht: (2026)
von: Penaloza, Emiliano, et al.
Veröffentlicht: (2026)
Unifying Group-Relative and Self-Distillation Policy Optimization via Sample Routing
von: Li, Gengsheng, et al.
Veröffentlicht: (2026)
von: Li, Gengsheng, et al.
Veröffentlicht: (2026)
PACED: Distillation and On-Policy Self-Distillation at the Frontier of Student Competence
von: Xu, Yuanda, et al.
Veröffentlicht: (2026)
von: Xu, Yuanda, et al.
Veröffentlicht: (2026)
Distilling Privileged Information for Dubins Traveling Salesman Problems with Neighborhoods
von: Shin, Min Kyu, et al.
Veröffentlicht: (2024)
von: Shin, Min Kyu, et al.
Veröffentlicht: (2024)
AVSD: Adaptive-View Self-Distillation by Balancing Consensus and Teacher-Specific Privileged Signals
von: Nguyen, Duy, et al.
Veröffentlicht: (2026)
von: Nguyen, Duy, et al.
Veröffentlicht: (2026)
TRACE: Distilling Where It Matters via Token-Routed Self On-Policy Alignment
von: Wang, Jiaxuan, et al.
Veröffentlicht: (2026)
von: Wang, Jiaxuan, et al.
Veröffentlicht: (2026)
Proximal Policy Distillation
von: Spigler, Giacomo
Veröffentlicht: (2024)
von: Spigler, Giacomo
Veröffentlicht: (2024)
Reinforcement Learning via Self-Distillation
von: Hübotter, Jonas, et al.
Veröffentlicht: (2026)
von: Hübotter, Jonas, et al.
Veröffentlicht: (2026)
ORPO-Distill: Mixed-Policy Preference Optimization for Cross-Architecture LLM Distillation
von: Singh, Aasheesh, et al.
Veröffentlicht: (2025)
von: Singh, Aasheesh, et al.
Veröffentlicht: (2025)
Extreme Region Policy Distillation
von: Chen, Changyu, et al.
Veröffentlicht: (2026)
von: Chen, Changyu, et al.
Veröffentlicht: (2026)
From Generic Correlation to Input-Specific Credit in On-Policy Self Distillation
von: Shen, Guobin, et al.
Veröffentlicht: (2026)
von: Shen, Guobin, et al.
Veröffentlicht: (2026)
TIP: Token Importance in On-Policy Distillation
von: Xu, Yuanda, et al.
Veröffentlicht: (2026)
von: Xu, Yuanda, et al.
Veröffentlicht: (2026)
Online Policy Distillation with Decision-Attention
von: Yu, Xinqiang, et al.
Veröffentlicht: (2024)
von: Yu, Xinqiang, et al.
Veröffentlicht: (2024)
When Are Teacher Tokens Reliable? Position-Weighted On-Policy Self-Distillation for Reasoning
von: Liu, Xiaogeng, et al.
Veröffentlicht: (2026)
von: Liu, Xiaogeng, et al.
Veröffentlicht: (2026)
VISTA: Validation-Informed Trajectory Adaptation via Self-Distillation
von: Corn, Eli, et al.
Veröffentlicht: (2026)
von: Corn, Eli, et al.
Veröffentlicht: (2026)
OGLS-SD: On-Policy Self-Distillation with Outcome-Guided Logit Steering for LLM Reasoning
von: Yang, Yuxiao, et al.
Veröffentlicht: (2026)
von: Yang, Yuxiao, et al.
Veröffentlicht: (2026)
OPD+: Rethinking the Advantage Design for On-Policy Distillation
von: Zhao, Hanyang, et al.
Veröffentlicht: (2026)
von: Zhao, Hanyang, et al.
Veröffentlicht: (2026)
Trust-Region Behavior Blending for On-Policy Distillation
von: Plyusov, Daniil, et al.
Veröffentlicht: (2026)
von: Plyusov, Daniil, et al.
Veröffentlicht: (2026)
Retrieval-Aware Distillation for Transformer-SSM Hybrids
von: Bick, Aviv, et al.
Veröffentlicht: (2026)
von: Bick, Aviv, et al.
Veröffentlicht: (2026)
The Mamba in the Llama: Distilling and Accelerating Hybrid Models
von: Wang, Junxiong, et al.
Veröffentlicht: (2024)
von: Wang, Junxiong, et al.
Veröffentlicht: (2024)
RLDG: Robotic Generalist Policy Distillation via Reinforcement Learning
von: Xu, Charles, et al.
Veröffentlicht: (2024)
von: Xu, Charles, et al.
Veröffentlicht: (2024)
Multilingual Safety Alignment via Self-Distillation
von: Qin, Ruiyang, et al.
Veröffentlicht: (2026)
von: Qin, Ruiyang, et al.
Veröffentlicht: (2026)
Self Distillation via Iterative Constructive Perturbations
von: Dave, Maheak, et al.
Veröffentlicht: (2025)
von: Dave, Maheak, et al.
Veröffentlicht: (2025)
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation
von: Liu, Junzhang, et al.
Veröffentlicht: (2024)
von: Liu, Junzhang, et al.
Veröffentlicht: (2024)
On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes
von: Agarwal, Rishabh, et al.
Veröffentlicht: (2023)
von: Agarwal, Rishabh, et al.
Veröffentlicht: (2023)
The Illusion of Certainty: Decoupling Capability and Calibration in On-Policy Distillation
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2026)
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2026)
ADWIN: Adaptive Windows for Horizon-Aware On-Policy Distillation
von: Liang, Kun, et al.
Veröffentlicht: (2026)
von: Liang, Kun, et al.
Veröffentlicht: (2026)
Stable On-Policy Distillation through Adaptive Target Reformulation
von: Jang, Ijun, et al.
Veröffentlicht: (2026)
von: Jang, Ijun, et al.
Veröffentlicht: (2026)
Interpretable Policy Distillation for Power Grid Topology Control
von: Dmitruka, Aleksandra, et al.
Veröffentlicht: (2026)
von: Dmitruka, Aleksandra, et al.
Veröffentlicht: (2026)
Evaluating Interpretable Reinforcement Learning by Distilling Policies into Programs
von: Kohler, Hector, et al.
Veröffentlicht: (2025)
von: Kohler, Hector, et al.
Veröffentlicht: (2025)
FlowDistill: Scalable Traffic Flow Prediction via Distillation from LLMs
von: Yu, Chenyang, et al.
Veröffentlicht: (2025)
von: Yu, Chenyang, et al.
Veröffentlicht: (2025)
FedSDR: Federated Self-Distillation with Rectification
von: Ren, Ziheng, et al.
Veröffentlicht: (2026)
von: Ren, Ziheng, et al.
Veröffentlicht: (2026)
Self-Distilled Disentangled Learning for Counterfactual Prediction
von: Li, Xinshu, et al.
Veröffentlicht: (2024)
von: Li, Xinshu, et al.
Veröffentlicht: (2024)
GNN's Uncertainty Quantification using Self-Distillation
von: Daneshvar, Hirad, et al.
Veröffentlicht: (2025)
von: Daneshvar, Hirad, et al.
Veröffentlicht: (2025)
OISD: On-Policy Internal Self-Distillation of Language Models
von: Liu, Xinyu, et al.
Veröffentlicht: (2026)
von: Liu, Xinyu, et al.
Veröffentlicht: (2026)
AlignDistil: Token-Level Language Model Alignment as Adaptive Policy Distillation
von: Zhang, Songming, et al.
Veröffentlicht: (2025)
von: Zhang, Songming, et al.
Veröffentlicht: (2025)
Variational Distillation of Diffusion Policies into Mixture of Experts
von: Zhou, Hongyi, et al.
Veröffentlicht: (2024)
von: Zhou, Hongyi, et al.
Veröffentlicht: (2024)
SHRED: Retain-Set-Free Unlearning via Self-Distillation with Logit Demotion
von: Hu, Zizhao, et al.
Veröffentlicht: (2026)
von: Hu, Zizhao, et al.
Veröffentlicht: (2026)
Learning from Partial Chain-of-Thought via Truncated-Reasoning Self-Distillation
von: Silvestri, Gianluigi, et al.
Veröffentlicht: (2026)
von: Silvestri, Gianluigi, et al.
Veröffentlicht: (2026)
Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization
von: Yu, Xin, et al.
Veröffentlicht: (2026)
von: Yu, Xin, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Privileged Information Distillation for Language Models
von: Penaloza, Emiliano, et al.
Veröffentlicht: (2026) -
Unifying Group-Relative and Self-Distillation Policy Optimization via Sample Routing
von: Li, Gengsheng, et al.
Veröffentlicht: (2026) -
PACED: Distillation and On-Policy Self-Distillation at the Frontier of Student Competence
von: Xu, Yuanda, et al.
Veröffentlicht: (2026) -
Distilling Privileged Information for Dubins Traveling Salesman Problems with Neighborhoods
von: Shin, Min Kyu, et al.
Veröffentlicht: (2024) -
AVSD: Adaptive-View Self-Distillation by Balancing Consensus and Teacher-Specific Privileged Signals
von: Nguyen, Duy, et al.
Veröffentlicht: (2026)