Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Yecheng, Han, Song, Cai, Hai |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Prune-OPD: Efficient and Reliable On-Policy Distillation for Long-Horizon Reasoning
von: Yang, Zhicheng, et al.
Veröffentlicht: (2026)
von: Yang, Zhicheng, et al.
Veröffentlicht: (2026)
OPD+: Rethinking the Advantage Design for On-Policy Distillation
von: Zhao, Hanyang, et al.
Veröffentlicht: (2026)
von: Zhao, Hanyang, et al.
Veröffentlicht: (2026)
DP-OPD: Differentially Private On-Policy Distillation for Language Models
von: Khadem, Fatemeh, et al.
Veröffentlicht: (2026)
von: Khadem, Fatemeh, et al.
Veröffentlicht: (2026)
$\boldsymbol{f}$-OPD: Stabilizing Long-Horizon On-Policy Distillation with Freshness-Aware Control
von: Chen, Xianwei, et al.
Veröffentlicht: (2026)
von: Chen, Xianwei, et al.
Veröffentlicht: (2026)
Decoupling KL and Trajectories: A Unified Perspective for SFT, DAgger, Offline RL, and OPD in LLM Distillation
von: Zhao, Anhao, et al.
Veröffentlicht: (2026)
von: Zhao, Anhao, et al.
Veröffentlicht: (2026)
MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate
von: Wang, Jianze, et al.
Veröffentlicht: (2026)
von: Wang, Jianze, et al.
Veröffentlicht: (2026)
SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models
von: Xiao, Guangxuan, et al.
Veröffentlicht: (2022)
von: Xiao, Guangxuan, et al.
Veröffentlicht: (2022)
GVPO: Group Variance Policy Optimization for Large Language Model Post-Training
von: Zhang, Kaichen, et al.
Veröffentlicht: (2025)
von: Zhang, Kaichen, et al.
Veröffentlicht: (2025)
Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models
von: Qu, Yun, et al.
Veröffentlicht: (2026)
von: Qu, Yun, et al.
Veröffentlicht: (2026)
Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation
von: Yuan, Qianhao, et al.
Veröffentlicht: (2026)
von: Yuan, Qianhao, et al.
Veröffentlicht: (2026)
Beyond GRPO and On-Policy Distillation: An Empirical Sparse-to-Dense Reward Principle for Language-Model Post-Training
von: Xu, Yuanda, et al.
Veröffentlicht: (2026)
von: Xu, Yuanda, et al.
Veröffentlicht: (2026)
ProteinOPD: Towards Effective and Efficient Preference Alignment for Protein Design
von: Zhang, Yulin, et al.
Veröffentlicht: (2026)
von: Zhang, Yulin, et al.
Veröffentlicht: (2026)
Offline Behavior Distillation
von: Lei, Shiye, et al.
Veröffentlicht: (2024)
von: Lei, Shiye, et al.
Veröffentlicht: (2024)
Beyond the Proxy: Trajectory-Distilled Guidance for Offline GFlowNet Training
von: Chen, Ruishuo, et al.
Veröffentlicht: (2025)
von: Chen, Ruishuo, et al.
Veröffentlicht: (2025)
Training Large Language Models to Reason via EM Policy Gradient
von: Xu, Tianbing
Veröffentlicht: (2025)
von: Xu, Tianbing
Veröffentlicht: (2025)
How to Compress KV Cache in RL Post-Training? Shadow Mask Distillation for Memory-Efficient Alignment
von: Zhu, Rui, et al.
Veröffentlicht: (2026)
von: Zhu, Rui, et al.
Veröffentlicht: (2026)
When Policies Cannot Be Retrained: A Unified Closed-Form View of Post-Training Steering in Offline Reinforcement Learning
von: Hossain, Elias, et al.
Veröffentlicht: (2026)
von: Hossain, Elias, et al.
Veröffentlicht: (2026)
IPD: Boosting Sequential Policy with Imaginary Planning Distillation in Offline Reinforcement Learning
von: Qin, Yihao, et al.
Veröffentlicht: (2026)
von: Qin, Yihao, et al.
Veröffentlicht: (2026)
Train at Moving Edge: Online-Verified Prompt Selection for Efficient RL Training of Large Reasoning Model
von: Wu, Jiahao, et al.
Veröffentlicht: (2026)
von: Wu, Jiahao, et al.
Veröffentlicht: (2026)
CROP: Conservative Reward for Model-based Offline Policy Optimization
von: Li, Hao, et al.
Veröffentlicht: (2023)
von: Li, Hao, et al.
Veröffentlicht: (2023)
Post-Training is About States, Not Tokens: A State Distribution View of SFT, RL, and On-Policy Distillation
von: Nie, Dong
Veröffentlicht: (2026)
von: Nie, Dong
Veröffentlicht: (2026)
Learning to Reason Efficiently with A* Post-Training
von: Opedal, Andreas, et al.
Veröffentlicht: (2026)
von: Opedal, Andreas, et al.
Veröffentlicht: (2026)
Agent Lightning: Train ANY AI Agents with Reinforcement Learning
von: Luo, Xufang, et al.
Veröffentlicht: (2025)
von: Luo, Xufang, et al.
Veröffentlicht: (2025)
Post-Training as Reweighting: A Stochastic View of Reasoning Trajectories in Language Models
von: Bu, Dake, et al.
Veröffentlicht: (2025)
von: Bu, Dake, et al.
Veröffentlicht: (2025)
Post-Training with Policy Gradients: Optimality and the Base Model Barrier
von: Mousavi-Hosseini, Alireza, et al.
Veröffentlicht: (2026)
von: Mousavi-Hosseini, Alireza, et al.
Veröffentlicht: (2026)
When is Offline Policy Selection Sample Efficient for Reinforcement Learning?
von: Liu, Vincent, et al.
Veröffentlicht: (2023)
von: Liu, Vincent, et al.
Veröffentlicht: (2023)
OM2P: Offline Multi-Agent Mean-Flow Policy
von: Li, Zhuoran, et al.
Veröffentlicht: (2025)
von: Li, Zhuoran, et al.
Veröffentlicht: (2025)
Dataset Distillation for Offline Reinforcement Learning
von: Light, Jonathan, et al.
Veröffentlicht: (2024)
von: Light, Jonathan, et al.
Veröffentlicht: (2024)
GLAI: GreenLightningAI for Accelerated Training through Knowledge Decoupling
von: Mestre, Jose I., et al.
Veröffentlicht: (2025)
von: Mestre, Jose I., et al.
Veröffentlicht: (2025)
Large Language Model Post-Training: A Unified View of Off-Policy and On-Policy Learning
von: Zhao, Shiwan, et al.
Veröffentlicht: (2026)
von: Zhao, Shiwan, et al.
Veröffentlicht: (2026)
Offline Learning and Forgetting for Reasoning with Large Language Models
von: Ni, Tianwei, et al.
Veröffentlicht: (2025)
von: Ni, Tianwei, et al.
Veröffentlicht: (2025)
ERPO: Token-Level Entropy-Regulated Policy Optimization for Large Reasoning Models
von: Yu, Song, et al.
Veröffentlicht: (2026)
von: Yu, Song, et al.
Veröffentlicht: (2026)
SDXL-Lightning: Progressive Adversarial Diffusion Distillation
von: Lin, Shanchuan, et al.
Veröffentlicht: (2024)
von: Lin, Shanchuan, et al.
Veröffentlicht: (2024)
Jet-Nemotron: Efficient Language Model with Post Neural Architecture Search
von: Gu, Yuxian, et al.
Veröffentlicht: (2025)
von: Gu, Yuxian, et al.
Veröffentlicht: (2025)
Preference Goal Tuning: Post-Training as Latent Control for Frozen Policies
von: Zhao, Guangyu, et al.
Veröffentlicht: (2024)
von: Zhao, Guangyu, et al.
Veröffentlicht: (2024)
KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning
von: Xu, Hongling, et al.
Veröffentlicht: (2025)
von: Xu, Hongling, et al.
Veröffentlicht: (2025)
One-Step Offline Distillation of Diffusion-based Models via Koopman Modeling
von: Berman, Nimrod, et al.
Veröffentlicht: (2025)
von: Berman, Nimrod, et al.
Veröffentlicht: (2025)
The Quest for Efficient Reasoning: A Data-Centric Benchmark to CoT Distillation
von: Zhang, Ruichen, et al.
Veröffentlicht: (2025)
von: Zhang, Ruichen, et al.
Veröffentlicht: (2025)
Hide to Guide: Learning via Semantic Masking
von: Liu, Ruitao, et al.
Veröffentlicht: (2026)
von: Liu, Ruitao, et al.
Veröffentlicht: (2026)
TED: Training-Free Experience Distillation for Multimodal Reasoning
von: Yuan, Shuozhi, et al.
Veröffentlicht: (2026)
von: Yuan, Shuozhi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Prune-OPD: Efficient and Reliable On-Policy Distillation for Long-Horizon Reasoning
von: Yang, Zhicheng, et al.
Veröffentlicht: (2026) -
OPD+: Rethinking the Advantage Design for On-Policy Distillation
von: Zhao, Hanyang, et al.
Veröffentlicht: (2026) -
DP-OPD: Differentially Private On-Policy Distillation for Language Models
von: Khadem, Fatemeh, et al.
Veröffentlicht: (2026) -
$\boldsymbol{f}$-OPD: Stabilizing Long-Horizon On-Policy Distillation with Freshness-Aware Control
von: Chen, Xianwei, et al.
Veröffentlicht: (2026) -
Decoupling KL and Trajectories: A Unified Perspective for SFT, DAgger, Offline RL, and OPD in LLM Distillation
von: Zhao, Anhao, et al.
Veröffentlicht: (2026)