Gespeichert in:
| Hauptverfasser: | Hou, Wenjin, Peng, Shangpin, Wang, Weinong, Ruan, Zheng, Zhang, Yue, Zhou, Zhenglin, Gao, Mingqi, Chen, Yifei, Wang, Kaiqi, Yang, Hongming, Zhang, Chengquan, Tian, Zhuotao, Hu, Han, Yang, Yi, Wu, Fei, Fan, Hehe |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2605.03677 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Uni-DPO: A Unified Paradigm for Dynamic Preference Optimization of LLMs
von: Peng, Shangpin, et al.
Veröffentlicht: (2025)
von: Peng, Shangpin, et al.
Veröffentlicht: (2025)
DiffusionOPD: A Unified Perspective of On-Policy Distillation in Diffusion Models
von: Li, Quanhao, et al.
Veröffentlicht: (2026)
von: Li, Quanhao, et al.
Veröffentlicht: (2026)
Mitigating Object Hallucinations via Sentence-Level Early Intervention
von: Peng, Shangpin, et al.
Veröffentlicht: (2025)
von: Peng, Shangpin, et al.
Veröffentlicht: (2025)
MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate
von: Wang, Jianze, et al.
Veröffentlicht: (2026)
von: Wang, Jianze, et al.
Veröffentlicht: (2026)
Prune-OPD: Efficient and Reliable On-Policy Distillation for Long-Horizon Reasoning
von: Yang, Zhicheng, et al.
Veröffentlicht: (2026)
von: Yang, Zhicheng, et al.
Veröffentlicht: (2026)
Draft-OPD: On-Policy Distillation for Speculative Draft Models
von: Lei, Haodi, et al.
Veröffentlicht: (2026)
von: Lei, Haodi, et al.
Veröffentlicht: (2026)
Incentivizing Generative Zero-Shot Learning via Outcome-Reward Reinforcement Learning with Visual Cues
von: Hou, Wenjin, et al.
Veröffentlicht: (2026)
von: Hou, Wenjin, et al.
Veröffentlicht: (2026)
OmniOPD: Logit-Free On-Policy Distillation via Speculative Verification
von: Zhou, Yuhang, et al.
Veröffentlicht: (2026)
von: Zhou, Yuhang, et al.
Veröffentlicht: (2026)
OPD+: Rethinking the Advantage Design for On-Policy Distillation
von: Zhao, Hanyang, et al.
Veröffentlicht: (2026)
von: Zhao, Hanyang, et al.
Veröffentlicht: (2026)
Adversarial Dual On-Policy Distillation from Expressive Teacher
von: Wan, Zhenglin, et al.
Veröffentlicht: (2026)
von: Wan, Zhenglin, et al.
Veröffentlicht: (2026)
Flow-OPD: On-Policy Distillation for Flow Matching Models
von: Fang, Zhen, et al.
Veröffentlicht: (2026)
von: Fang, Zhen, et al.
Veröffentlicht: (2026)
Safe, or Simply Incapable? Rethinking Safety Evaluation for Phone-Use Agents
von: Tang, Zhengyang, et al.
Veröffentlicht: (2026)
von: Tang, Zhengyang, et al.
Veröffentlicht: (2026)
DP-OPD: Differentially Private On-Policy Distillation for Language Models
von: Khadem, Fatemeh, et al.
Veröffentlicht: (2026)
von: Khadem, Fatemeh, et al.
Veröffentlicht: (2026)
$\boldsymbol{f}$-OPD: Stabilizing Long-Horizon On-Policy Distillation with Freshness-Aware Control
von: Chen, Xianwei, et al.
Veröffentlicht: (2026)
von: Chen, Xianwei, et al.
Veröffentlicht: (2026)
Deepfake Detection Generalization with Diffusion Noise
von: Qi, Hongyuan, et al.
Veröffentlicht: (2026)
von: Qi, Hongyuan, et al.
Veröffentlicht: (2026)
Decoupling KL and Trajectories: A Unified Perspective for SFT, DAgger, Offline RL, and OPD in LLM Distillation
von: Zhao, Anhao, et al.
Veröffentlicht: (2026)
von: Zhao, Anhao, et al.
Veröffentlicht: (2026)
HeadStudio: Text to Animatable Head Avatars with 3D Gaussian Splatting
von: Zhou, Zhenglin, et al.
Veröffentlicht: (2024)
von: Zhou, Zhenglin, et al.
Veröffentlicht: (2024)
EDGE-OPD: Internalizing Privileged Context with Evidence Guided On-Policy Distillation
von: Lazaridis, Aristotelis, et al.
Veröffentlicht: (2026)
von: Lazaridis, Aristotelis, et al.
Veröffentlicht: (2026)
Prompt-Aware Controllable Shadow Removal
von: Chen, Kerui, et al.
Veröffentlicht: (2025)
von: Chen, Kerui, et al.
Veröffentlicht: (2025)
ZeroMamba: Exploring Visual State Space Model for Zero-Shot Learning
von: Hou, Wenjin, et al.
Veröffentlicht: (2024)
von: Hou, Wenjin, et al.
Veröffentlicht: (2024)
Translution: Unifying Self-attention and Convolution for Adaptive and Relative Modeling
von: Fan, Hehe, et al.
Veröffentlicht: (2025)
von: Fan, Hehe, et al.
Veröffentlicht: (2025)
Depictor : Topic‐Guided Opinion Summarization for Product Reviews With Dual‐Perspective Topic Modeling
von: Yanyue Zhang, et al.
Veröffentlicht: (2026)
von: Yanyue Zhang, et al.
Veröffentlicht: (2026)
X-OPD: Cross-Modal On-Policy Distillation for Capability Alignment in Speech LLMs
von: Cao, Di, et al.
Veröffentlicht: (2026)
von: Cao, Di, et al.
Veröffentlicht: (2026)
PhoneWorld: Scaling Phone-Use Agent Environments
von: Tang, Zhengyang, et al.
Veröffentlicht: (2026)
von: Tang, Zhengyang, et al.
Veröffentlicht: (2026)
Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
von: Li, Yaxuan, et al.
Veröffentlicht: (2026)
von: Li, Yaxuan, et al.
Veröffentlicht: (2026)
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts
von: Zhang, Yue, et al.
Veröffentlicht: (2025)
von: Zhang, Yue, et al.
Veröffentlicht: (2025)
Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation
von: Wu, Yecheng, et al.
Veröffentlicht: (2026)
von: Wu, Yecheng, et al.
Veröffentlicht: (2026)
VLA-OPD: Bridging Offline SFT and Online RL for Vision-Language-Action Models via On-Policy Distillation
von: Zhong, Zhide, et al.
Veröffentlicht: (2026)
von: Zhong, Zhide, et al.
Veröffentlicht: (2026)
Uni-Instruct: One-step Diffusion Model through Unified Diffusion Divergence Instruction
von: Wang, Yifei, et al.
Veröffentlicht: (2025)
von: Wang, Yifei, et al.
Veröffentlicht: (2025)
UniGraspTransformer: Simplified Policy Distillation for Scalable Dexterous Robotic Grasping
von: Wang, Wenbo, et al.
Veröffentlicht: (2024)
von: Wang, Wenbo, et al.
Veröffentlicht: (2024)
Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation
von: Yuan, Qianhao, et al.
Veröffentlicht: (2026)
von: Yuan, Qianhao, et al.
Veröffentlicht: (2026)
Zero-1-to-A: Zero-Shot One Image to Animatable Head Avatars Using Video Diffusion
von: Zhou, Zhenglin, et al.
Veröffentlicht: (2025)
von: Zhou, Zhenglin, et al.
Veröffentlicht: (2025)
Enhancing MLLM Spatial Understanding via Active 3D Scene Exploration for Multi-Perspective Reasoning
von: Chen, Jiahua, et al.
Veröffentlicht: (2026)
von: Chen, Jiahua, et al.
Veröffentlicht: (2026)
CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies
von: Du, Fan, et al.
Veröffentlicht: (2026)
von: Du, Fan, et al.
Veröffentlicht: (2026)
ITS3D: Inference-Time Scaling for Text-Guided 3D Diffusion Models
von: Zhou, Zhenglin, et al.
Veröffentlicht: (2025)
von: Zhou, Zhenglin, et al.
Veröffentlicht: (2025)
DreamDPO: Aligning Text-to-3D Generation with Human Preferences via Direct Preference Optimization
von: Zhou, Zhenglin, et al.
Veröffentlicht: (2025)
von: Zhou, Zhenglin, et al.
Veröffentlicht: (2025)
ANO: A Principled Approach to Robust Policy Optimization
von: Zhang, Yiheng, et al.
Veröffentlicht: (2026)
von: Zhang, Yiheng, et al.
Veröffentlicht: (2026)
Unified Language-driven Zero-shot Domain Adaptation
von: Yang, Senqiao, et al.
Veröffentlicht: (2024)
von: Yang, Senqiao, et al.
Veröffentlicht: (2024)
Seeing Is Believing? A Benchmark for Multimodal Large Language Models on Visual Illusions and Anomalies
von: Hou, Wenjin, et al.
Veröffentlicht: (2026)
von: Hou, Wenjin, et al.
Veröffentlicht: (2026)
SedarEval: Automated Evaluation using Self-Adaptive Rubrics
von: Fan, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Fan, Zhiyuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Uni-DPO: A Unified Paradigm for Dynamic Preference Optimization of LLMs
von: Peng, Shangpin, et al.
Veröffentlicht: (2025) -
DiffusionOPD: A Unified Perspective of On-Policy Distillation in Diffusion Models
von: Li, Quanhao, et al.
Veröffentlicht: (2026) -
Mitigating Object Hallucinations via Sentence-Level Early Intervention
von: Peng, Shangpin, et al.
Veröffentlicht: (2025) -
MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate
von: Wang, Jianze, et al.
Veröffentlicht: (2026) -
Prune-OPD: Efficient and Reliable On-Policy Distillation for Long-Horizon Reasoning
von: Yang, Zhicheng, et al.
Veröffentlicht: (2026)