Learning to Reason under Off-Policy Guidance
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yan, Jianhao, Li, Yafu, Hu, Zican, Wang, Zhi, Cui, Ganqu, Qu, Xiaoye, Cheng, Yu, Zhang, Yue |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ExGRPO: Learning to Reason from Experience
von: Zhan, Runzhe, et al.
Veröffentlicht: (2025)
von: Zhan, Runzhe, et al.
Veröffentlicht: (2025)
Diversity-Incentivized Exploration for Versatile Reasoning
von: Hu, Zican, et al.
Veröffentlicht: (2025)
von: Hu, Zican, et al.
Veröffentlicht: (2025)
Advancing Multimodal Reasoning: From Optimized Cold Start to Staged Reinforcement Learning
von: Chen, Shuang, et al.
Veröffentlicht: (2025)
von: Chen, Shuang, et al.
Veröffentlicht: (2025)
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models
von: Fu, Tingchen, et al.
Veröffentlicht: (2025)
von: Fu, Tingchen, et al.
Veröffentlicht: (2025)
Rethinking Entropy Regularization in Large Reasoning Models
von: Jiang, Yuxian, et al.
Veröffentlicht: (2025)
von: Jiang, Yuxian, et al.
Veröffentlicht: (2025)
A Survey of Reinforcement Learning for Large Reasoning Models
von: Zhang, Kaiyan, et al.
Veröffentlicht: (2025)
von: Zhang, Kaiyan, et al.
Veröffentlicht: (2025)
The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
von: Cui, Ganqu, et al.
Veröffentlicht: (2025)
von: Cui, Ganqu, et al.
Veröffentlicht: (2025)
SEE: Continual Fine-tuning with Sequential Ensemble of Experts
von: Wang, Zhilin, et al.
Veröffentlicht: (2025)
von: Wang, Zhilin, et al.
Veröffentlicht: (2025)
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid
von: Sun, Weigao, et al.
Veröffentlicht: (2025)
von: Sun, Weigao, et al.
Veröffentlicht: (2025)
PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models
von: Song, Mingyang, et al.
Veröffentlicht: (2025)
von: Song, Mingyang, et al.
Veröffentlicht: (2025)
Large Language Model Post-Training: A Unified View of Off-Policy and On-Policy Learning
von: Zhao, Shiwan, et al.
Veröffentlicht: (2026)
von: Zhao, Shiwan, et al.
Veröffentlicht: (2026)
Potential and Challenges of Model Editing for Social Debiasing
von: Yan, Jianhao, et al.
Veröffentlicht: (2024)
von: Yan, Jianhao, et al.
Veröffentlicht: (2024)
Twin-Merging: Dynamic Integration of Modular Expertise in Model Merging
von: Lu, Zhenyi, et al.
Veröffentlicht: (2024)
von: Lu, Zhenyi, et al.
Veröffentlicht: (2024)
BAPO: Stabilizing Off-Policy Reinforcement Learning for LLMs via Balanced Policy Optimization with Adaptive Clipping
von: Xi, Zhiheng, et al.
Veröffentlicht: (2025)
von: Xi, Zhiheng, et al.
Veröffentlicht: (2025)
POPE: Learning to Reason on Hard Problems via Privileged On-Policy Exploration
von: Qu, Yuxiao, et al.
Veröffentlicht: (2026)
von: Qu, Yuxiao, et al.
Veröffentlicht: (2026)
Slow-Fast Policy Optimization: Reposition-Before-Update for LLM Reasoning
von: Wang, Ziyan, et al.
Veröffentlicht: (2025)
von: Wang, Ziyan, et al.
Veröffentlicht: (2025)
From Drafts to Answers: Unlocking LLM Potential via Aggregation Fine-Tuning
von: Li, Yafu, et al.
Veröffentlicht: (2025)
von: Li, Yafu, et al.
Veröffentlicht: (2025)
Scalable Efficient Training of Large Language Models with Low-dimensional Projected Attention
von: Lv, Xingtai, et al.
Veröffentlicht: (2024)
von: Lv, Xingtai, et al.
Veröffentlicht: (2024)
Nested-ReFT: Efficient Reinforcement Learning for Large Language Model Fine-Tuning via Off-Policy Rollouts
von: Heuillet, Maxime, et al.
Veröffentlicht: (2025)
von: Heuillet, Maxime, et al.
Veröffentlicht: (2025)
On Giant's Shoulders: Effortless Weak to Strong by Dynamic Logits Fusion
von: Fan, Chenghao, et al.
Veröffentlicht: (2024)
von: Fan, Chenghao, et al.
Veröffentlicht: (2024)
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models
von: Nath, Vaskar, et al.
Veröffentlicht: (2025)
von: Nath, Vaskar, et al.
Veröffentlicht: (2025)
Dual-Uncertainty Guided Policy Learning for Multimodal Reasoning
von: Liu, Rui, et al.
Veröffentlicht: (2025)
von: Liu, Rui, et al.
Veröffentlicht: (2025)
CLIPO: Contrastive Learning in Policy Optimization Generalizes RLVR
von: Cui, Sijia, et al.
Veröffentlicht: (2026)
von: Cui, Sijia, et al.
Veröffentlicht: (2026)
Off-Policy Corrected Reward Modeling for Reinforcement Learning from Human Feedback
von: Ackermann, Johannes, et al.
Veröffentlicht: (2025)
von: Ackermann, Johannes, et al.
Veröffentlicht: (2025)
Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance
von: Yan, Kai, et al.
Veröffentlicht: (2026)
von: Yan, Kai, et al.
Veröffentlicht: (2026)
DOTS: Learning to Reason Dynamically in LLMs via Optimal Reasoning Trajectories Search
von: Yue, Murong, et al.
Veröffentlicht: (2024)
von: Yue, Murong, et al.
Veröffentlicht: (2024)
Linear-MoE: Linear Sequence Modeling Meets Mixture-of-Experts
von: Sun, Weigao, et al.
Veröffentlicht: (2025)
von: Sun, Weigao, et al.
Veröffentlicht: (2025)
Learning to Reason for Hallucination Span Detection
von: Su, Hsuan, et al.
Veröffentlicht: (2025)
von: Su, Hsuan, et al.
Veröffentlicht: (2025)
Advancing LLM Reasoning Generalists with Preference Trees
von: Yuan, Lifan, et al.
Veröffentlicht: (2024)
von: Yuan, Lifan, et al.
Veröffentlicht: (2024)
Divide and Conquer: Grounding LLMs as Efficient Decision-Making Agents via Offline Hierarchical Reinforcement Learning
von: Hu, Zican, et al.
Veröffentlicht: (2025)
von: Hu, Zican, et al.
Veröffentlicht: (2025)
FlowRL: Matching Reward Distributions for LLM Reasoning
von: Zhu, Xuekai, et al.
Veröffentlicht: (2025)
von: Zhu, Xuekai, et al.
Veröffentlicht: (2025)
Generative Adversarial Reasoner: Enhancing LLM Reasoning with Adversarial Reinforcement Learning
von: Liu, Qihao, et al.
Veröffentlicht: (2025)
von: Liu, Qihao, et al.
Veröffentlicht: (2025)
Klear-Reasoner: Advancing Reasoning Capability via Gradient-Preserving Clipping Policy Optimization
von: Su, Zhenpeng, et al.
Veröffentlicht: (2025)
von: Su, Zhenpeng, et al.
Veröffentlicht: (2025)
ReasonFlux: Hierarchical LLM Reasoning via Scaling Thought Templates
von: Yang, Ling, et al.
Veröffentlicht: (2025)
von: Yang, Ling, et al.
Veröffentlicht: (2025)
Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models
von: Noukhovitch, Michael, et al.
Veröffentlicht: (2024)
von: Noukhovitch, Michael, et al.
Veröffentlicht: (2024)
StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason
von: Zhang, Kaiyi, et al.
Veröffentlicht: (2025)
von: Zhang, Kaiyi, et al.
Veröffentlicht: (2025)
RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning
von: Mao, Yixiu, et al.
Veröffentlicht: (2026)
von: Mao, Yixiu, et al.
Veröffentlicht: (2026)
Look Inward to Explore Outward: Learning Temperature Policy from LLM Internal States via Hierarchical RL
von: Zhou, Yixiao, et al.
Veröffentlicht: (2026)
von: Zhou, Yixiao, et al.
Veröffentlicht: (2026)
Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM Reasoning
von: Zhang, Xichen, et al.
Veröffentlicht: (2025)
von: Zhang, Xichen, et al.
Veröffentlicht: (2025)
Guidance is All You Need: Temperature-Guided Reasoning in Large Language Models
von: Gomaa, Eyad, et al.
Veröffentlicht: (2024)
von: Gomaa, Eyad, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
ExGRPO: Learning to Reason from Experience
von: Zhan, Runzhe, et al.
Veröffentlicht: (2025) -
Diversity-Incentivized Exploration for Versatile Reasoning
von: Hu, Zican, et al.
Veröffentlicht: (2025) -
Advancing Multimodal Reasoning: From Optimized Cold Start to Staged Reinforcement Learning
von: Chen, Shuang, et al.
Veröffentlicht: (2025) -
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models
von: Fu, Tingchen, et al.
Veröffentlicht: (2025) -
Rethinking Entropy Regularization in Large Reasoning Models
von: Jiang, Yuxian, et al.
Veröffentlicht: (2025)