DeCoRL: Decoupling Reasoning Chains via Parallel Sub-Step Generation and Cascaded Reinforcement for Interpretable and Scalable RLHF
Fuente:
arXiv
Saved in:
| Main Authors: | Gao, Ziyuan, Liang, Di, Wu, Xianjie, Morel, Philippe, Peng, Minlong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Structural Reward Model: Enhancing Interpretability, Efficiency, and Scalability in Reward Modeling
by: Liu, Xiaoyu, et al.
Published: (2025)
by: Liu, Xiaoyu, et al.
Published: (2025)
Prompt-Aware Adaptive Elastic Weight Consolidation for Continual Learning in Medical Vision-Language Models
by: Gao, Ziyuan, et al.
Published: (2025)
by: Gao, Ziyuan, et al.
Published: (2025)
DPL: Spatial-Conditioned Diffusion Prototype Enhancement for One-Shot Medical Segmentation
by: Gao, Ziyuan, et al.
Published: (2025)
by: Gao, Ziyuan, et al.
Published: (2025)
Not All Parameters Are Created Equal: Smart Isolation Boosts Fine-Tuning Performance
by: Wang, Yao, et al.
Published: (2025)
by: Wang, Yao, et al.
Published: (2025)
Boosting Deductive Reasoning with Step Signals In RLHF
by: Li, Jialian, et al.
Published: (2024)
by: Li, Jialian, et al.
Published: (2024)
Reinforcement Learning Enhanced Multi-hop Reasoning for Temporal Knowledge Question Answering
by: Wen, Wuzhenghong, et al.
Published: (2026)
by: Wen, Wuzhenghong, et al.
Published: (2026)
Reason Only When Needed: Efficient Generative Reward Modeling via Model-Internal Uncertainty
by: Xue, Chao, et al.
Published: (2026)
by: Xue, Chao, et al.
Published: (2026)
DeReason: A Difficulty-Aware Curriculum Improves Decoupled SFT-then-RL Training for General Reasoning
by: Hu, Hanxu, et al.
Published: (2026)
by: Hu, Hanxu, et al.
Published: (2026)
Decoupled Behavioral Cloning for Scalable Inductive Generalization in RL from Specifications
by: Subramanian, Vignesh, et al.
Published: (2026)
by: Subramanian, Vignesh, et al.
Published: (2026)
HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical Decoupling
by: Liu, Xianjie, et al.
Published: (2025)
by: Liu, Xianjie, et al.
Published: (2025)
Divide-and-Conquer CoT: RL for Reducing Latency via Parallel Reasoning
by: Mahankali, Arvind, et al.
Published: (2026)
by: Mahankali, Arvind, et al.
Published: (2026)
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
by: Zhang, Yi-Fan, et al.
Published: (2025)
by: Zhang, Yi-Fan, et al.
Published: (2025)
Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step
by: Guo, Ziyu, et al.
Published: (2025)
by: Guo, Ziyu, et al.
Published: (2025)
Parameter Importance is Not Static: Evolving Parameter Isolation for Supervised Fine-Tuning
by: Lin, Zekai, et al.
Published: (2026)
by: Lin, Zekai, et al.
Published: (2026)
RL Tango: Reinforcing Generator and Verifier Together for Language Reasoning
by: Zha, Kaiwen, et al.
Published: (2025)
by: Zha, Kaiwen, et al.
Published: (2025)
An Adaptive Placement and Parallelism Framework for Accelerating RLHF Training
by: Xiao, Youshao, et al.
Published: (2023)
by: Xiao, Youshao, et al.
Published: (2023)
Proxy-RLHF: Decoupling Generation and Alignment in Large Language Model with Proxy
by: Zhu, Yu, et al.
Published: (2024)
by: Zhu, Yu, et al.
Published: (2024)
Latent Chain-of-Thought as Planning: Decoupling Reasoning from Verbalization
by: Wang, Jiecong, et al.
Published: (2026)
by: Wang, Jiecong, et al.
Published: (2026)
OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
by: Hu, Jian, et al.
Published: (2024)
by: Hu, Jian, et al.
Published: (2024)
VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning
by: Lu, Guanxing, et al.
Published: (2025)
by: Lu, Guanxing, et al.
Published: (2025)
Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use
by: Goldie, Anna, et al.
Published: (2025)
by: Goldie, Anna, et al.
Published: (2025)
Nemotron-Cascade: Scaling Cascaded Reinforcement Learning for General-Purpose Reasoning Models
by: Wang, Boxin, et al.
Published: (2025)
by: Wang, Boxin, et al.
Published: (2025)
A Hopfieldian View-based Interpretation for Chain-of-Thought Reasoning
by: Hu, Lijie, et al.
Published: (2024)
by: Hu, Lijie, et al.
Published: (2024)
DRAFT-RL: Multi-Agent Chain-of-Draft Reasoning for Reinforcement Learning-Enhanced LLMs
by: Li, Yuanhao, et al.
Published: (2025)
by: Li, Yuanhao, et al.
Published: (2025)
TaxonRL: Reinforcement Learning with Intermediate Rewards for Interpretable Fine-Grained Visual Reasoning
by: von Klinski, Maximilian, et al.
Published: (2026)
by: von Klinski, Maximilian, et al.
Published: (2026)
DPO Meets PPO: Reinforced Token Optimization for RLHF
by: Zhong, Han, et al.
Published: (2024)
by: Zhong, Han, et al.
Published: (2024)
Decoupling Intrinsic Molecular Efficacy from Platform Effects: An Interpretable Machine Learning Framework for Unbiased Perovskite Passivator Discovery
by: Zhang, Jing, et al.
Published: (2026)
by: Zhang, Jing, et al.
Published: (2026)
Decoupled Prioritized Resampling for Offline RL
by: Yue, Yang, et al.
Published: (2023)
by: Yue, Yang, et al.
Published: (2023)
One Framework to Rule Them All: Unifying RL-Based and RL-Free Methods in RLHF
by: Cai, Xin
Published: (2025)
by: Cai, Xin
Published: (2025)
Imaging of a Case for Primary Pulmonary Dedifferentiated Liposarcoma
by: Tongxin Yang, et al.
Published: (2025)
by: Tongxin Yang, et al.
Published: (2025)
Reasoning-Aligned Perception Decoupling for Scalable Multi-modal Reasoning
by: Gou, Yunhao, et al.
Published: (2025)
by: Gou, Yunhao, et al.
Published: (2025)
MarsRL: Advancing Multi-Agent Reasoning System via Reinforcement Learning with Agentic Pipeline Parallelism
by: Liu, Shulin, et al.
Published: (2025)
by: Liu, Shulin, et al.
Published: (2025)
MoRL: Reinforced Reasoning for Unified Motion Understanding and Generation
by: Wang, Hongpeng, et al.
Published: (2026)
by: Wang, Hongpeng, et al.
Published: (2026)
Accelerating Long-Tail Generation in Synchronous RLHF Training via Adaptive Tensor Parallelism
by: Zhao, Long, et al.
Published: (2026)
by: Zhao, Long, et al.
Published: (2026)
Prototype-Guided Classification Sub-Task Decoupling Framework: Enhancing Generalization and Interpretability for Multivariate Time Series
by: Song, Xianhao, et al.
Published: (2026)
by: Song, Xianhao, et al.
Published: (2026)
Step Guided Reasoning: Improving Mathematical Reasoning using Guidance Generation and Step Reasoning
by: Cao, Lang, et al.
Published: (2024)
by: Cao, Lang, et al.
Published: (2024)
Chain-of-Probe: Examining the Necessity and Accuracy of CoT Step-by-Step
by: Wang, Zezhong, et al.
Published: (2024)
by: Wang, Zezhong, et al.
Published: (2024)
Do Latent-CoT Models Think Step-by-Step? A Mechanistic Study on Sequential Reasoning Tasks
by: Liang, Jia, et al.
Published: (2026)
by: Liang, Jia, et al.
Published: (2026)
CAPS: Cascaded Adaptive Pairwise Selection for Efficient Parallel Reasoning
by: Lin, Fangzhou, et al.
Published: (2026)
by: Lin, Fangzhou, et al.
Published: (2026)
LLaVA-CoT: Let Vision Language Models Reason Step-by-Step
by: Xu, Guowei, et al.
Published: (2024)
by: Xu, Guowei, et al.
Published: (2024)
Similar Items
-
Structural Reward Model: Enhancing Interpretability, Efficiency, and Scalability in Reward Modeling
by: Liu, Xiaoyu, et al.
Published: (2025) -
Prompt-Aware Adaptive Elastic Weight Consolidation for Continual Learning in Medical Vision-Language Models
by: Gao, Ziyuan, et al.
Published: (2025) -
DPL: Spatial-Conditioned Diffusion Prototype Enhancement for One-Shot Medical Segmentation
by: Gao, Ziyuan, et al.
Published: (2025) -
Not All Parameters Are Created Equal: Smart Isolation Boosts Fine-Tuning Performance
by: Wang, Yao, et al.
Published: (2025) -
Boosting Deductive Reasoning with Step Signals In RLHF
by: Li, Jialian, et al.
Published: (2024)