Saved in:
| Main Authors: | Yang, Han, Wu, Mingyan, He, Bailan, Cao, Zeyu, Yan, Sikuan, Lin, Kevin Qinghong, Ding, Zifeng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2605.08776 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TCP: a Benchmark for Temporal Constraint-Based Planning
by: Ding, Zifeng, et al.
Published: (2025)
by: Ding, Zifeng, et al.
Published: (2025)
PERFT: Parameter-Efficient Routed Fine-Tuning for Mixture-of-Expert Model
by: Liu, Yilun, et al.
Published: (2024)
by: Liu, Yilun, et al.
Published: (2024)
EigentSearch-Q+: Enhancing Deep Research Agents with Structured Reasoning Tools
by: Zhang, Boer, et al.
Published: (2026)
by: Zhang, Boer, et al.
Published: (2026)
Temporal Fact Reasoning over Hyper-Relational Knowledge Graphs
by: Ding, Zifeng, et al.
Published: (2023)
by: Ding, Zifeng, et al.
Published: (2023)
Paper2Video: Automatic Video Generation from Scientific Papers
by: Zhu, Zeyu, et al.
Published: (2025)
by: Zhu, Zeyu, et al.
Published: (2025)
Surgical Post-Training: Proximal On-Policy Distillation for Reasoning with Knowledge Retention
by: Lin, Wenye, et al.
Published: (2026)
by: Lin, Wenye, et al.
Published: (2026)
Consistency Policy: Accelerated Visuomotor Policies via Consistency Distillation
by: Prasad, Aaditya, et al.
Published: (2024)
by: Prasad, Aaditya, et al.
Published: (2024)
Routing-Free Mixture-of-Experts
by: Liu, Yilun, et al.
Published: (2026)
by: Liu, Yilun, et al.
Published: (2026)
Think or Not? Selective Reasoning via Reinforcement Learning for Vision-Language Models
by: Wang, Jiaqi, et al.
Published: (2025)
by: Wang, Jiaqi, et al.
Published: (2025)
Structural Rationale Distillation via Reasoning Space Compression
by: Yang, Jialin, et al.
Published: (2026)
by: Yang, Jialin, et al.
Published: (2026)
Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation
by: Wu, Yecheng, et al.
Published: (2026)
by: Wu, Yecheng, et al.
Published: (2026)
VideoMind: A Chain-of-LoRA Agent for Temporal-Grounded Video Reasoning
by: Liu, Ye, et al.
Published: (2025)
by: Liu, Ye, et al.
Published: (2025)
Respecting Self-Uncertainty in On-Policy Self-Distillation for Efficient LLM Reasoning
by: Ke, Junlong, et al.
Published: (2026)
by: Ke, Junlong, et al.
Published: (2026)
Revisiting 3D LLM Benchmarks: Are We Really Testing 3D Capabilities?
by: Jin, Jiahe, et al.
Published: (2025)
by: Jin, Jiahe, et al.
Published: (2025)
HDPO: Hybrid Distillation Policy Optimization via Privileged Self-Distillation
by: Ding, Ken
Published: (2026)
by: Ding, Ken
Published: (2026)
Can Knowledge Graphs Make Large Language Models More Trustworthy? An Empirical Study Over Open-ended Question Answering
by: Sui, Yuan, et al.
Published: (2024)
by: Sui, Yuan, et al.
Published: (2024)
OPD+: Rethinking the Advantage Design for On-Policy Distillation
by: Zhao, Hanyang, et al.
Published: (2026)
by: Zhao, Hanyang, et al.
Published: (2026)
ORPO-Distill: Mixed-Policy Preference Optimization for Cross-Architecture LLM Distillation
by: Singh, Aasheesh, et al.
Published: (2025)
by: Singh, Aasheesh, et al.
Published: (2025)
DRPO: Efficient Reasoning via Decoupled Reward Policy Optimization
by: Li, Gang, et al.
Published: (2025)
by: Li, Gang, et al.
Published: (2025)
Extreme Region Policy Distillation
by: Chen, Changyu, et al.
Published: (2026)
by: Chen, Changyu, et al.
Published: (2026)
Prune-OPD: Efficient and Reliable On-Policy Distillation for Long-Horizon Reasoning
by: Yang, Zhicheng, et al.
Published: (2026)
by: Yang, Zhicheng, et al.
Published: (2026)
Mixed Distillation Helps Smaller Language Model Better Reasoning
by: Li, Chenglin, et al.
Published: (2023)
by: Li, Chenglin, et al.
Published: (2023)
ShowUI-$π$: Flow-based Generative Models as GUI Dexterous Hands
by: Hu, Siyuan, et al.
Published: (2025)
by: Hu, Siyuan, et al.
Published: (2025)
Clinical-R1: Empowering Large Language Models for Faithful and Comprehensive Reasoning with Clinical Objective Relative Policy Optimization
by: Gu, Boyang, et al.
Published: (2025)
by: Gu, Boyang, et al.
Published: (2025)
SceneCode: Executable World Programs for Editable Indoor Scenes with Articulated Objects
by: Wang, Puyi, et al.
Published: (2026)
by: Wang, Puyi, et al.
Published: (2026)
OGLS-SD: On-Policy Self-Distillation with Outcome-Guided Logit Steering for LLM Reasoning
by: Yang, Yuxiao, et al.
Published: (2026)
by: Yang, Yuxiao, et al.
Published: (2026)
Paper2Poster: Towards Multimodal Poster Automation from Scientific Papers
by: Pang, Wei, et al.
Published: (2025)
by: Pang, Wei, et al.
Published: (2025)
KDMOS:Knowledge Distillation for Motion Segmentation
by: Cao, Chunyu, et al.
Published: (2025)
by: Cao, Chunyu, et al.
Published: (2025)
PACED: Distillation and On-Policy Self-Distillation at the Frontier of Student Competence
by: Xu, Yuanda, et al.
Published: (2026)
by: Xu, Yuanda, et al.
Published: (2026)
MCPO: Mastery-Consolidated Policy Optimization for Large Reasoning Models
by: Liao, Zhaokang, et al.
Published: (2026)
by: Liao, Zhaokang, et al.
Published: (2026)
LLM-Augmented Digital Twin for Policy Evaluation in Short-Video Platforms
by: Zhang, Haoting, et al.
Published: (2026)
by: Zhang, Haoting, et al.
Published: (2026)
OPSDL: On-Policy Self-Distillation for Long-Context Language Models
by: Zhang, Xinsen, et al.
Published: (2026)
by: Zhang, Xinsen, et al.
Published: (2026)
ManiCM: Real-time 3D Diffusion Policy via Consistency Model for Robotic Manipulation
by: Lu, Guanxing, et al.
Published: (2024)
by: Lu, Guanxing, et al.
Published: (2024)
Data-Efficient On-Policy Distillation for Automatic Speech Recognition
by: Lin, Yu, et al.
Published: (2026)
by: Lin, Yu, et al.
Published: (2026)
EchoRL: Reinforcement Learning via Rollout Echoing
by: Bi, Jinhe, et al.
Published: (2026)
by: Bi, Jinhe, et al.
Published: (2026)
Flow-OPD: On-Policy Distillation for Flow Matching Models
by: Fang, Zhen, et al.
Published: (2026)
by: Fang, Zhen, et al.
Published: (2026)
TIP: Token Importance in On-Policy Distillation
by: Xu, Yuanda, et al.
Published: (2026)
by: Xu, Yuanda, et al.
Published: (2026)
LRAS: Advanced Legal Reasoning with Agentic Search
by: Zhou, Yujin, et al.
Published: (2026)
by: Zhou, Yujin, et al.
Published: (2026)
WeatherReasonSeg: A Benchmark for Weather-Aware Reasoning Segmentation in Visual Language Models
by: Du, Wanjun, et al.
Published: (2026)
by: Du, Wanjun, et al.
Published: (2026)
Code2Video: A Code-centric Paradigm for Educational Video Generation
by: Chen, Yanzhe, et al.
Published: (2025)
by: Chen, Yanzhe, et al.
Published: (2025)
Similar Items
-
TCP: a Benchmark for Temporal Constraint-Based Planning
by: Ding, Zifeng, et al.
Published: (2025) -
PERFT: Parameter-Efficient Routed Fine-Tuning for Mixture-of-Expert Model
by: Liu, Yilun, et al.
Published: (2024) -
EigentSearch-Q+: Enhancing Deep Research Agents with Structured Reasoning Tools
by: Zhang, Boer, et al.
Published: (2026) -
Temporal Fact Reasoning over Hyper-Relational Knowledge Graphs
by: Ding, Zifeng, et al.
Published: (2023) -
Paper2Video: Automatic Video Generation from Scientific Papers
by: Zhu, Zeyu, et al.
Published: (2025)