Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe
Fuente:
arXiv
Salvato in:
| Autori principali: | Hou, Wenjin, Peng, Shangpin, Wang, Weinong, Ruan, Zheng, Zhang, Yue, Zhou, Zhenglin, Gao, Mingqi, Chen, Yifei, Wang, Kaiqi, Yang, Hongming, Zhang, Chengquan, Tian, Zhuotao, Hu, Han, Yang, Yi, Wu, Fei, Fan, Hehe |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Uni-DPO: A Unified Paradigm for Dynamic Preference Optimization of LLMs
di: Peng, Shangpin, et al.
Pubblicazione: (2025)
di: Peng, Shangpin, et al.
Pubblicazione: (2025)
DiffusionOPD: A Unified Perspective of On-Policy Distillation in Diffusion Models
di: Li, Quanhao, et al.
Pubblicazione: (2026)
di: Li, Quanhao, et al.
Pubblicazione: (2026)
Mitigating Object Hallucinations via Sentence-Level Early Intervention
di: Peng, Shangpin, et al.
Pubblicazione: (2025)
di: Peng, Shangpin, et al.
Pubblicazione: (2025)
MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate
di: Wang, Jianze, et al.
Pubblicazione: (2026)
di: Wang, Jianze, et al.
Pubblicazione: (2026)
Prune-OPD: Efficient and Reliable On-Policy Distillation for Long-Horizon Reasoning
di: Yang, Zhicheng, et al.
Pubblicazione: (2026)
di: Yang, Zhicheng, et al.
Pubblicazione: (2026)
Draft-OPD: On-Policy Distillation for Speculative Draft Models
di: Lei, Haodi, et al.
Pubblicazione: (2026)
di: Lei, Haodi, et al.
Pubblicazione: (2026)
OmniOPD: Logit-Free On-Policy Distillation via Speculative Verification
di: Zhou, Yuhang, et al.
Pubblicazione: (2026)
di: Zhou, Yuhang, et al.
Pubblicazione: (2026)
OPD+: Rethinking the Advantage Design for On-Policy Distillation
di: Zhao, Hanyang, et al.
Pubblicazione: (2026)
di: Zhao, Hanyang, et al.
Pubblicazione: (2026)
Adversarial Dual On-Policy Distillation from Expressive Teacher
di: Wan, Zhenglin, et al.
Pubblicazione: (2026)
di: Wan, Zhenglin, et al.
Pubblicazione: (2026)
Incentivizing Generative Zero-Shot Learning via Outcome-Reward Reinforcement Learning with Visual Cues
di: Hou, Wenjin, et al.
Pubblicazione: (2026)
di: Hou, Wenjin, et al.
Pubblicazione: (2026)
Flow-OPD: On-Policy Distillation for Flow Matching Models
di: Fang, Zhen, et al.
Pubblicazione: (2026)
di: Fang, Zhen, et al.
Pubblicazione: (2026)
$\boldsymbol{f}$-OPD: Stabilizing Long-Horizon On-Policy Distillation with Freshness-Aware Control
di: Chen, Xianwei, et al.
Pubblicazione: (2026)
di: Chen, Xianwei, et al.
Pubblicazione: (2026)
DP-OPD: Differentially Private On-Policy Distillation for Language Models
di: Khadem, Fatemeh, et al.
Pubblicazione: (2026)
di: Khadem, Fatemeh, et al.
Pubblicazione: (2026)
Decoupling KL and Trajectories: A Unified Perspective for SFT, DAgger, Offline RL, and OPD in LLM Distillation
di: Zhao, Anhao, et al.
Pubblicazione: (2026)
di: Zhao, Anhao, et al.
Pubblicazione: (2026)
EDGE-OPD: Internalizing Privileged Context with Evidence Guided On-Policy Distillation
di: Lazaridis, Aristotelis, et al.
Pubblicazione: (2026)
di: Lazaridis, Aristotelis, et al.
Pubblicazione: (2026)
Deepfake Detection Generalization with Diffusion Noise
di: Qi, Hongyuan, et al.
Pubblicazione: (2026)
di: Qi, Hongyuan, et al.
Pubblicazione: (2026)
X-OPD: Cross-Modal On-Policy Distillation for Capability Alignment in Speech LLMs
di: Cao, Di, et al.
Pubblicazione: (2026)
di: Cao, Di, et al.
Pubblicazione: (2026)
HeadStudio: Text to Animatable Head Avatars with 3D Gaussian Splatting
di: Zhou, Zhenglin, et al.
Pubblicazione: (2024)
di: Zhou, Zhenglin, et al.
Pubblicazione: (2024)
Safe, or Simply Incapable? Rethinking Safety Evaluation for Phone-Use Agents
di: Tang, Zhengyang, et al.
Pubblicazione: (2026)
di: Tang, Zhengyang, et al.
Pubblicazione: (2026)
Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
di: Li, Yaxuan, et al.
Pubblicazione: (2026)
di: Li, Yaxuan, et al.
Pubblicazione: (2026)
Depictor : Topic‐Guided Opinion Summarization for Product Reviews With Dual‐Perspective Topic Modeling
di: Yanyue Zhang, et al.
Pubblicazione: (2026)
di: Yanyue Zhang, et al.
Pubblicazione: (2026)
Translution: Unifying Self-attention and Convolution for Adaptive and Relative Modeling
di: Fan, Hehe, et al.
Pubblicazione: (2025)
di: Fan, Hehe, et al.
Pubblicazione: (2025)
Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation
di: Wu, Yecheng, et al.
Pubblicazione: (2026)
di: Wu, Yecheng, et al.
Pubblicazione: (2026)
VLA-OPD: Bridging Offline SFT and Online RL for Vision-Language-Action Models via On-Policy Distillation
di: Zhong, Zhide, et al.
Pubblicazione: (2026)
di: Zhong, Zhide, et al.
Pubblicazione: (2026)
Prompt-Aware Controllable Shadow Removal
di: Chen, Kerui, et al.
Pubblicazione: (2025)
di: Chen, Kerui, et al.
Pubblicazione: (2025)
ZeroMamba: Exploring Visual State Space Model for Zero-Shot Learning
di: Hou, Wenjin, et al.
Pubblicazione: (2024)
di: Hou, Wenjin, et al.
Pubblicazione: (2024)
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts
di: Zhang, Yue, et al.
Pubblicazione: (2025)
di: Zhang, Yue, et al.
Pubblicazione: (2025)
Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation
di: Yuan, Qianhao, et al.
Pubblicazione: (2026)
di: Yuan, Qianhao, et al.
Pubblicazione: (2026)
UniGraspTransformer: Simplified Policy Distillation for Scalable Dexterous Robotic Grasping
di: Wang, Wenbo, et al.
Pubblicazione: (2024)
di: Wang, Wenbo, et al.
Pubblicazione: (2024)
Uni-Instruct: One-step Diffusion Model through Unified Diffusion Divergence Instruction
di: Wang, Yifei, et al.
Pubblicazione: (2025)
di: Wang, Yifei, et al.
Pubblicazione: (2025)
PhoneWorld: Scaling Phone-Use Agent Environments
di: Tang, Zhengyang, et al.
Pubblicazione: (2026)
di: Tang, Zhengyang, et al.
Pubblicazione: (2026)
Enhancing MLLM Spatial Understanding via Active 3D Scene Exploration for Multi-Perspective Reasoning
di: Chen, Jiahua, et al.
Pubblicazione: (2026)
di: Chen, Jiahua, et al.
Pubblicazione: (2026)
Zero-1-to-A: Zero-Shot One Image to Animatable Head Avatars Using Video Diffusion
di: Zhou, Zhenglin, et al.
Pubblicazione: (2025)
di: Zhou, Zhenglin, et al.
Pubblicazione: (2025)
CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies
di: Du, Fan, et al.
Pubblicazione: (2026)
di: Du, Fan, et al.
Pubblicazione: (2026)
CoDistill-GRPO: A Co-Distillation Recipe for Efficient Group Relative Policy Optimization
di: Kwon, Soo Min, et al.
Pubblicazione: (2026)
di: Kwon, Soo Min, et al.
Pubblicazione: (2026)
Unified Language-driven Zero-shot Domain Adaptation
di: Yang, Senqiao, et al.
Pubblicazione: (2024)
di: Yang, Senqiao, et al.
Pubblicazione: (2024)
ANO: A Principled Approach to Robust Policy Optimization
di: Zhang, Yiheng, et al.
Pubblicazione: (2026)
di: Zhang, Yiheng, et al.
Pubblicazione: (2026)
Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation
di: Li, Jiaze, et al.
Pubblicazione: (2026)
di: Li, Jiaze, et al.
Pubblicazione: (2026)
Uni-Renderer: Unifying Rendering and Inverse Rendering Via Dual Stream Diffusion
di: Chen, Zhifei, et al.
Pubblicazione: (2024)
di: Chen, Zhifei, et al.
Pubblicazione: (2024)
VividDreamer: Invariant Score Distillation For Hyper-Realistic Text-to-3D Generation
di: Zhuo, Wenjie, et al.
Pubblicazione: (2024)
di: Zhuo, Wenjie, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Uni-DPO: A Unified Paradigm for Dynamic Preference Optimization of LLMs
di: Peng, Shangpin, et al.
Pubblicazione: (2025) -
DiffusionOPD: A Unified Perspective of On-Policy Distillation in Diffusion Models
di: Li, Quanhao, et al.
Pubblicazione: (2026) -
Mitigating Object Hallucinations via Sentence-Level Early Intervention
di: Peng, Shangpin, et al.
Pubblicazione: (2025) -
MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate
di: Wang, Jianze, et al.
Pubblicazione: (2026) -
Prune-OPD: Efficient and Reliable On-Policy Distillation for Long-Horizon Reasoning
di: Yang, Zhicheng, et al.
Pubblicazione: (2026)