The Many Faces of On-Policy Distillation: Pitfalls, Mechanisms, and Fixes
Fuente:
arXiv
Saved in:
| Main Authors: | Zhu, Siqi, Ye, Xuyan, Lu, Hongyu, Shi, Weiye, Liu, Ge |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes
by: Fu, Yuqian, et al.
Published: (2026)
by: Fu, Yuqian, et al.
Published: (2026)
Distilling Generative-Discriminative Representations for Very Low-Resolution Face Recognition
by: Zhang, Junzheng, et al.
Published: (2024)
by: Zhang, Junzheng, et al.
Published: (2024)
Low-Resolution Face Recognition via Adaptable Instance-Relation Distillation
by: Shi, Ruixin, et al.
Published: (2024)
by: Shi, Ruixin, et al.
Published: (2024)
Hybrid Policy Distillation for LLMs
by: Zhu, Wenhong, et al.
Published: (2026)
by: Zhu, Wenhong, et al.
Published: (2026)
DARC: Decoupled Asymmetric Reasoning Curriculum for LLM Evolution
by: Fan, Shengda, et al.
Published: (2026)
by: Fan, Shengda, et al.
Published: (2026)
CA*: Addressing Evaluation Pitfalls in Computation-Aware Latency for Simultaneous Speech Translation
by: Xu, Xi, et al.
Published: (2024)
by: Xu, Xi, et al.
Published: (2024)
SafeSteer: Localized On-Policy Distillation for Efficient Safety Alignment
by: Li, Hao, et al.
Published: (2026)
by: Li, Hao, et al.
Published: (2026)
Absolute Policy Optimization
by: Zhao, Weiye, et al.
Published: (2023)
by: Zhao, Weiye, et al.
Published: (2023)
Your Teacher Can't Help You Here: Combating Supervision Fidelity Decay in On-Policy Distillation
by: Liu, Yanjiang, et al.
Published: (2026)
by: Liu, Yanjiang, et al.
Published: (2026)
Offline Semantic Guidance for Efficient Vision-Language-Action Policy Distillation
by: Shi, Jin, et al.
Published: (2026)
by: Shi, Jin, et al.
Published: (2026)
It is Too Many Options: Pitfalls of Multiple-Choice Questions in Generative AI and Medical Education
by: Singh, Shrutika, et al.
Published: (2025)
by: Singh, Shrutika, et al.
Published: (2025)
Efficient Low-Resolution Face Recognition via Bridge Distillation
by: Ge, Shiming, et al.
Published: (2024)
by: Ge, Shiming, et al.
Published: (2024)
Variational Distillation of Diffusion Policies into Mixture of Experts
by: Zhou, Hongyi, et al.
Published: (2024)
by: Zhou, Hongyi, et al.
Published: (2024)
Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
by: Li, Yaxuan, et al.
Published: (2026)
by: Li, Yaxuan, et al.
Published: (2026)
Look Through Masks: Towards Masked Face Recognition with De-Occlusion Distillation
by: Li, Chenyu, et al.
Published: (2024)
by: Li, Chenyu, et al.
Published: (2024)
Black-Box On-Policy Distillation of Large Language Models
by: Ye, Tianzhu, et al.
Published: (2025)
by: Ye, Tianzhu, et al.
Published: (2025)
The Illusion of Certainty: Decoupling Capability and Calibration in On-Policy Distillation
by: Zhang, Jiaxin, et al.
Published: (2026)
by: Zhang, Jiaxin, et al.
Published: (2026)
LLMs Know More Than Words: A Genre Study with Syntax, Metaphor & Phonetics
by: Shi, Weiye, et al.
Published: (2025)
by: Shi, Weiye, et al.
Published: (2025)
EDGE-OPD: Internalizing Privileged Context with Evidence Guided On-Policy Distillation
by: Lazaridis, Aristotelis, et al.
Published: (2026)
by: Lazaridis, Aristotelis, et al.
Published: (2026)
Same Evidence, Different Answers: Canonical-Context On-Policy Distillation for Multi-Turn Language Models
by: Lin, Zizhuo, et al.
Published: (2026)
by: Lin, Zizhuo, et al.
Published: (2026)
Reasoning Compression with Mixed-Policy Distillation
by: Yang, Han, et al.
Published: (2026)
by: Yang, Han, et al.
Published: (2026)
X-OPD: Cross-Modal On-Policy Distillation for Capability Alignment in Speech LLMs
by: Cao, Di, et al.
Published: (2026)
by: Cao, Di, et al.
Published: (2026)
Proximal Policy Distillation
by: Spigler, Giacomo
Published: (2024)
by: Spigler, Giacomo
Published: (2024)
Uncertainty-Aware 3D Emotional Talking Face Synthesis with Emotion Prior Distillation
by: Shen, Nanhan, et al.
Published: (2026)
by: Shen, Nanhan, et al.
Published: (2026)
Implicit Safe Set Algorithm for Provably Safe Reinforcement Learning
by: Zhao, Weiye, et al.
Published: (2024)
by: Zhao, Weiye, et al.
Published: (2024)
Who Prices Cognitive Labor in the Age of Agents? Compute-Anchored Wages
by: Zhu, Siqi
Published: (2026)
by: Zhu, Siqi
Published: (2026)
Agentic AI Systems Should Be Designed as Marginal Token Allocators
by: Zhu, Siqi
Published: (2026)
by: Zhu, Siqi
Published: (2026)
Accelerating Diffusion Models with One-to-Many Knowledge Distillation
by: Zhang, Linfeng, et al.
Published: (2024)
by: Zhang, Linfeng, et al.
Published: (2024)
GRExplainer: A Universal Explanation Method for Temporal Graph Neural Networks
by: Li, Xuyan, et al.
Published: (2025)
by: Li, Xuyan, et al.
Published: (2025)
Consistency Policy: Accelerated Visuomotor Policies via Consistency Distillation
by: Prasad, Aaditya, et al.
Published: (2024)
by: Prasad, Aaditya, et al.
Published: (2024)
Extreme Region Policy Distillation
by: Chen, Changyu, et al.
Published: (2026)
by: Chen, Changyu, et al.
Published: (2026)
ManiCM: Real-time 3D Diffusion Policy via Consistency Model for Robotic Manipulation
by: Lu, Guanxing, et al.
Published: (2024)
by: Lu, Guanxing, et al.
Published: (2024)
The Pitfalls of KV Cache Compression
by: Chen, Alex, et al.
Published: (2025)
by: Chen, Alex, et al.
Published: (2025)
Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation
by: Yuan, Qianhao, et al.
Published: (2026)
by: Yuan, Qianhao, et al.
Published: (2026)
PACED: Distillation and On-Policy Self-Distillation at the Frontier of Student Competence
by: Xu, Yuanda, et al.
Published: (2026)
by: Xu, Yuanda, et al.
Published: (2026)
Bandwidth-Efficient and Privacy-Preserving Edge-Cloud Many-to-Many Speech Translation
by: Du, Yexing, et al.
Published: (2026)
by: Du, Yexing, et al.
Published: (2026)
Data-Efficient On-Policy Distillation for Automatic Speech Recognition
by: Lin, Yu, et al.
Published: (2026)
by: Lin, Yu, et al.
Published: (2026)
TCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous Agents
by: Wang, Jiaqi, et al.
Published: (2026)
by: Wang, Jiaqi, et al.
Published: (2026)
One-Step Flow Policy: Self-Distillation for Fast Visuomotor Policies
by: Li, Shaolong, et al.
Published: (2026)
by: Li, Shaolong, et al.
Published: (2026)
ADWIN: Adaptive Windows for Horizon-Aware On-Policy Distillation
by: Liang, Kun, et al.
Published: (2026)
by: Liang, Kun, et al.
Published: (2026)
Similar Items
-
Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes
by: Fu, Yuqian, et al.
Published: (2026) -
Distilling Generative-Discriminative Representations for Very Low-Resolution Face Recognition
by: Zhang, Junzheng, et al.
Published: (2024) -
Low-Resolution Face Recognition via Adaptable Instance-Relation Distillation
by: Shi, Ruixin, et al.
Published: (2024) -
Hybrid Policy Distillation for LLMs
by: Zhu, Wenhong, et al.
Published: (2026) -
DARC: Decoupled Asymmetric Reasoning Curriculum for LLM Evolution
by: Fan, Shengda, et al.
Published: (2026)