DeCoRL: Decoupling Reasoning Chains via Parallel Sub-Step Generation and Cascaded Reinforcement for Interpretable and Scalable RLHF
Fuente:
arXiv
Salvato in:
| Autori principali: | Gao, Ziyuan, Liang, Di, Wu, Xianjie, Morel, Philippe, Peng, Minlong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Structural Reward Model: Enhancing Interpretability, Efficiency, and Scalability in Reward Modeling
di: Liu, Xiaoyu, et al.
Pubblicazione: (2025)
di: Liu, Xiaoyu, et al.
Pubblicazione: (2025)
Prompt-Aware Adaptive Elastic Weight Consolidation for Continual Learning in Medical Vision-Language Models
di: Gao, Ziyuan, et al.
Pubblicazione: (2025)
di: Gao, Ziyuan, et al.
Pubblicazione: (2025)
DPL: Spatial-Conditioned Diffusion Prototype Enhancement for One-Shot Medical Segmentation
di: Gao, Ziyuan, et al.
Pubblicazione: (2025)
di: Gao, Ziyuan, et al.
Pubblicazione: (2025)
Not All Parameters Are Created Equal: Smart Isolation Boosts Fine-Tuning Performance
di: Wang, Yao, et al.
Pubblicazione: (2025)
di: Wang, Yao, et al.
Pubblicazione: (2025)
Boosting Deductive Reasoning with Step Signals In RLHF
di: Li, Jialian, et al.
Pubblicazione: (2024)
di: Li, Jialian, et al.
Pubblicazione: (2024)
Reinforcement Learning Enhanced Multi-hop Reasoning for Temporal Knowledge Question Answering
di: Wen, Wuzhenghong, et al.
Pubblicazione: (2026)
di: Wen, Wuzhenghong, et al.
Pubblicazione: (2026)
Reason Only When Needed: Efficient Generative Reward Modeling via Model-Internal Uncertainty
di: Xue, Chao, et al.
Pubblicazione: (2026)
di: Xue, Chao, et al.
Pubblicazione: (2026)
DeReason: A Difficulty-Aware Curriculum Improves Decoupled SFT-then-RL Training for General Reasoning
di: Hu, Hanxu, et al.
Pubblicazione: (2026)
di: Hu, Hanxu, et al.
Pubblicazione: (2026)
Decoupled Behavioral Cloning for Scalable Inductive Generalization in RL from Specifications
di: Subramanian, Vignesh, et al.
Pubblicazione: (2026)
di: Subramanian, Vignesh, et al.
Pubblicazione: (2026)
HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical Decoupling
di: Liu, Xianjie, et al.
Pubblicazione: (2025)
di: Liu, Xianjie, et al.
Pubblicazione: (2025)
Divide-and-Conquer CoT: RL for Reducing Latency via Parallel Reasoning
di: Mahankali, Arvind, et al.
Pubblicazione: (2026)
di: Mahankali, Arvind, et al.
Pubblicazione: (2026)
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
di: Zhang, Yi-Fan, et al.
Pubblicazione: (2025)
di: Zhang, Yi-Fan, et al.
Pubblicazione: (2025)
Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step
di: Guo, Ziyu, et al.
Pubblicazione: (2025)
di: Guo, Ziyu, et al.
Pubblicazione: (2025)
Parameter Importance is Not Static: Evolving Parameter Isolation for Supervised Fine-Tuning
di: Lin, Zekai, et al.
Pubblicazione: (2026)
di: Lin, Zekai, et al.
Pubblicazione: (2026)
RL Tango: Reinforcing Generator and Verifier Together for Language Reasoning
di: Zha, Kaiwen, et al.
Pubblicazione: (2025)
di: Zha, Kaiwen, et al.
Pubblicazione: (2025)
An Adaptive Placement and Parallelism Framework for Accelerating RLHF Training
di: Xiao, Youshao, et al.
Pubblicazione: (2023)
di: Xiao, Youshao, et al.
Pubblicazione: (2023)
Proxy-RLHF: Decoupling Generation and Alignment in Large Language Model with Proxy
di: Zhu, Yu, et al.
Pubblicazione: (2024)
di: Zhu, Yu, et al.
Pubblicazione: (2024)
Latent Chain-of-Thought as Planning: Decoupling Reasoning from Verbalization
di: Wang, Jiecong, et al.
Pubblicazione: (2026)
di: Wang, Jiecong, et al.
Pubblicazione: (2026)
OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
di: Hu, Jian, et al.
Pubblicazione: (2024)
di: Hu, Jian, et al.
Pubblicazione: (2024)
VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning
di: Lu, Guanxing, et al.
Pubblicazione: (2025)
di: Lu, Guanxing, et al.
Pubblicazione: (2025)
Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use
di: Goldie, Anna, et al.
Pubblicazione: (2025)
di: Goldie, Anna, et al.
Pubblicazione: (2025)
Nemotron-Cascade: Scaling Cascaded Reinforcement Learning for General-Purpose Reasoning Models
di: Wang, Boxin, et al.
Pubblicazione: (2025)
di: Wang, Boxin, et al.
Pubblicazione: (2025)
A Hopfieldian View-based Interpretation for Chain-of-Thought Reasoning
di: Hu, Lijie, et al.
Pubblicazione: (2024)
di: Hu, Lijie, et al.
Pubblicazione: (2024)
DRAFT-RL: Multi-Agent Chain-of-Draft Reasoning for Reinforcement Learning-Enhanced LLMs
di: Li, Yuanhao, et al.
Pubblicazione: (2025)
di: Li, Yuanhao, et al.
Pubblicazione: (2025)
TaxonRL: Reinforcement Learning with Intermediate Rewards for Interpretable Fine-Grained Visual Reasoning
di: von Klinski, Maximilian, et al.
Pubblicazione: (2026)
di: von Klinski, Maximilian, et al.
Pubblicazione: (2026)
DPO Meets PPO: Reinforced Token Optimization for RLHF
di: Zhong, Han, et al.
Pubblicazione: (2024)
di: Zhong, Han, et al.
Pubblicazione: (2024)
Decoupling Intrinsic Molecular Efficacy from Platform Effects: An Interpretable Machine Learning Framework for Unbiased Perovskite Passivator Discovery
di: Zhang, Jing, et al.
Pubblicazione: (2026)
di: Zhang, Jing, et al.
Pubblicazione: (2026)
Decoupled Prioritized Resampling for Offline RL
di: Yue, Yang, et al.
Pubblicazione: (2023)
di: Yue, Yang, et al.
Pubblicazione: (2023)
One Framework to Rule Them All: Unifying RL-Based and RL-Free Methods in RLHF
di: Cai, Xin
Pubblicazione: (2025)
di: Cai, Xin
Pubblicazione: (2025)
Imaging of a Case for Primary Pulmonary Dedifferentiated Liposarcoma
di: Tongxin Yang, et al.
Pubblicazione: (2025)
di: Tongxin Yang, et al.
Pubblicazione: (2025)
Reasoning-Aligned Perception Decoupling for Scalable Multi-modal Reasoning
di: Gou, Yunhao, et al.
Pubblicazione: (2025)
di: Gou, Yunhao, et al.
Pubblicazione: (2025)
MarsRL: Advancing Multi-Agent Reasoning System via Reinforcement Learning with Agentic Pipeline Parallelism
di: Liu, Shulin, et al.
Pubblicazione: (2025)
di: Liu, Shulin, et al.
Pubblicazione: (2025)
MoRL: Reinforced Reasoning for Unified Motion Understanding and Generation
di: Wang, Hongpeng, et al.
Pubblicazione: (2026)
di: Wang, Hongpeng, et al.
Pubblicazione: (2026)
Accelerating Long-Tail Generation in Synchronous RLHF Training via Adaptive Tensor Parallelism
di: Zhao, Long, et al.
Pubblicazione: (2026)
di: Zhao, Long, et al.
Pubblicazione: (2026)
Prototype-Guided Classification Sub-Task Decoupling Framework: Enhancing Generalization and Interpretability for Multivariate Time Series
di: Song, Xianhao, et al.
Pubblicazione: (2026)
di: Song, Xianhao, et al.
Pubblicazione: (2026)
Step Guided Reasoning: Improving Mathematical Reasoning using Guidance Generation and Step Reasoning
di: Cao, Lang, et al.
Pubblicazione: (2024)
di: Cao, Lang, et al.
Pubblicazione: (2024)
Chain-of-Probe: Examining the Necessity and Accuracy of CoT Step-by-Step
di: Wang, Zezhong, et al.
Pubblicazione: (2024)
di: Wang, Zezhong, et al.
Pubblicazione: (2024)
Do Latent-CoT Models Think Step-by-Step? A Mechanistic Study on Sequential Reasoning Tasks
di: Liang, Jia, et al.
Pubblicazione: (2026)
di: Liang, Jia, et al.
Pubblicazione: (2026)
CAPS: Cascaded Adaptive Pairwise Selection for Efficient Parallel Reasoning
di: Lin, Fangzhou, et al.
Pubblicazione: (2026)
di: Lin, Fangzhou, et al.
Pubblicazione: (2026)
LLaVA-CoT: Let Vision Language Models Reason Step-by-Step
di: Xu, Guowei, et al.
Pubblicazione: (2024)
di: Xu, Guowei, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Structural Reward Model: Enhancing Interpretability, Efficiency, and Scalability in Reward Modeling
di: Liu, Xiaoyu, et al.
Pubblicazione: (2025) -
Prompt-Aware Adaptive Elastic Weight Consolidation for Continual Learning in Medical Vision-Language Models
di: Gao, Ziyuan, et al.
Pubblicazione: (2025) -
DPL: Spatial-Conditioned Diffusion Prototype Enhancement for One-Shot Medical Segmentation
di: Gao, Ziyuan, et al.
Pubblicazione: (2025) -
Not All Parameters Are Created Equal: Smart Isolation Boosts Fine-Tuning Performance
di: Wang, Yao, et al.
Pubblicazione: (2025) -
Boosting Deductive Reasoning with Step Signals In RLHF
di: Li, Jialian, et al.
Pubblicazione: (2024)