Boosting Deductive Reasoning with Step Signals In RLHF
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Jialian, Zhang, Yipin, Shen, Wei, Yan, Yuzi, Xie, Jian, Yan, Dong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reward-Robust RLHF in LLMs
by: Yan, Yuzi, et al.
Published: (2024)
by: Yan, Yuzi, et al.
Published: (2024)
3D-Properties: Identifying Challenges in DPO and Charting a Path Forward
by: Yan, Yuzi, et al.
Published: (2024)
by: Yan, Yuzi, et al.
Published: (2024)
Exploring the LLM Journey from Cognition to Expression with Linear Representations
by: Yan, Yuzi, et al.
Published: (2024)
by: Yan, Yuzi, et al.
Published: (2024)
Policy Filtration for RLHF to Mitigate Noise in Reward Models
by: Zhang, Chuheng, et al.
Published: (2024)
by: Zhang, Chuheng, et al.
Published: (2024)
RLHF Workflow: From Reward Modeling to Online RLHF
by: Dong, Hanze, et al.
Published: (2024)
by: Dong, Hanze, et al.
Published: (2024)
OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
by: Hu, Jian, et al.
Published: (2024)
by: Hu, Jian, et al.
Published: (2024)
The Role of Deductive and Inductive Reasoning in Large Language Models
by: Cai, Chengkun, et al.
Published: (2024)
by: Cai, Chengkun, et al.
Published: (2024)
The Signal is in the Steps: Local Scoring for Reasoning Data Selection
by: Just, Hoang Anh, et al.
Published: (2025)
by: Just, Hoang Anh, et al.
Published: (2025)
LLM Reasoning as Trajectories: Step-Specific Representation Geometry and Correctness Signals
by: Sun, Lihao, et al.
Published: (2026)
by: Sun, Lihao, et al.
Published: (2026)
Conditional Equivalence of DPO and RLHF: Implicit Assumption, Failure Modes, and Provable Alignment
by: Yang, Zhiqin, et al.
Published: (2026)
by: Yang, Zhiqin, et al.
Published: (2026)
A Shared Low-Rank Adaptation Approach to Personalized RLHF
by: Liu, Renpu, et al.
Published: (2025)
by: Liu, Renpu, et al.
Published: (2025)
ROCM: RLHF on consistency models
by: Shekhar, Shivanshu, et al.
Published: (2025)
by: Shekhar, Shivanshu, et al.
Published: (2025)
CPL: Critical Plan Step Learning Boosts LLM Generalization in Reasoning Tasks
by: Wang, Tianlong, et al.
Published: (2024)
by: Wang, Tianlong, et al.
Published: (2024)
When LLM Meets Time Series: Can LLMs Perform Multi-Step Time Series Reasoning and Inference
by: Ye, Wen, et al.
Published: (2025)
by: Ye, Wen, et al.
Published: (2025)
FlashMol: High-Quality Molecule Generation in as Few as Four Steps
by: Wei, Xinyuan, et al.
Published: (2026)
by: Wei, Xinyuan, et al.
Published: (2026)
Multi-Step Deductive Reasoning Over Natural Language: An Empirical Study on Out-of-Distribution Generalisation
by: Bao, Qiming, et al.
Published: (2022)
by: Bao, Qiming, et al.
Published: (2022)
Greedy Sampling Is Provably Efficient for RLHF
by: Wu, Di, et al.
Published: (2025)
by: Wu, Di, et al.
Published: (2025)
Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
by: Xiong, Wei, et al.
Published: (2023)
by: Xiong, Wei, et al.
Published: (2023)
AdaptiveStep: Automatically Dividing Reasoning Step through Model Confidence
by: Liu, Yuliang, et al.
Published: (2025)
by: Liu, Yuliang, et al.
Published: (2025)
What Are Step-Level Reward Models Rewarding? Counterintuitive Findings from MCTS-Boosted Mathematical Reasoning
by: Ma, Yiran, et al.
Published: (2024)
by: Ma, Yiran, et al.
Published: (2024)
RLHF and IIA: Perverse Incentives
by: Xu, Wanqiao, et al.
Published: (2023)
by: Xu, Wanqiao, et al.
Published: (2023)
Uncertainty-aware Reward Model: Teaching Reward Models to Know What is Unknown
by: Lou, Xingzhou, et al.
Published: (2024)
by: Lou, Xingzhou, et al.
Published: (2024)
Deciphering Scientific Reasoning Steps from Outcome Data for Molecule Optimization
by: Liu, Zequn, et al.
Published: (2026)
by: Liu, Zequn, et al.
Published: (2026)
The Hidden Link Between RLHF and Contrastive Learning
by: Lv, Xufei, et al.
Published: (2025)
by: Lv, Xufei, et al.
Published: (2025)
DeduCE: Deductive Consistency as a Framework to Evaluate LLM Reasoning
by: Pandey, Atharva, et al.
Published: (2025)
by: Pandey, Atharva, et al.
Published: (2025)
StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason
by: Zhang, Kaiyi, et al.
Published: (2025)
by: Zhang, Kaiyi, et al.
Published: (2025)
SSPO: Self-traced Step-wise Preference Optimization for Process Supervision and Reasoning Compression
by: Xu, Yuyang, et al.
Published: (2025)
by: Xu, Yuyang, et al.
Published: (2025)
Reward Generalization in RLHF: A Topological Perspective
by: Qiu, Tianyi, et al.
Published: (2024)
by: Qiu, Tianyi, et al.
Published: (2024)
BoostStep: Boosting mathematical capability of Large Language Models via improved single-step reasoning
by: Zhang, Beichen, et al.
Published: (2025)
by: Zhang, Beichen, et al.
Published: (2025)
KL-regularization Itself is Differentially Private in Bandits and RLHF
by: Zhang, Yizhou, et al.
Published: (2025)
by: Zhang, Yizhou, et al.
Published: (2025)
Towards Reward Fairness in RLHF: From a Resource Allocation Perspective
by: Ouyang, Sheng, et al.
Published: (2025)
by: Ouyang, Sheng, et al.
Published: (2025)
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF
by: Zhao, Kangwen, et al.
Published: (2025)
by: Zhao, Kangwen, et al.
Published: (2025)
Distributionally Robust Token Optimization in RLHF
by: Jin, Yeping, et al.
Published: (2026)
by: Jin, Yeping, et al.
Published: (2026)
Boosting Maximum Entropy Reinforcement Learning via One-Step Flow Matching
by: Li, Zeqiao, et al.
Published: (2026)
by: Li, Zeqiao, et al.
Published: (2026)
SplitZip: Ultra Fast Lossless KV Compression for Disaggregated LLM Serving
by: Guo, Yipin, et al.
Published: (2026)
by: Guo, Yipin, et al.
Published: (2026)
How to Evaluate Reward Models for RLHF
by: Frick, Evan, et al.
Published: (2024)
by: Frick, Evan, et al.
Published: (2024)
GPTQT: Quantize Large Language Models Twice to Push the Efficiency
by: Guo, Yipin, et al.
Published: (2024)
by: Guo, Yipin, et al.
Published: (2024)
ShiftAddViT: Mixture of Multiplication Primitives Towards Efficient Vision Transformer
by: You, Haoran, et al.
Published: (2023)
by: You, Haoran, et al.
Published: (2023)
Supervised Reinforcement Learning: From Expert Trajectories to Step-wise Reasoning
by: Deng, Yihe, et al.
Published: (2025)
by: Deng, Yihe, et al.
Published: (2025)
RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation
by: Park, Chanwoo, et al.
Published: (2024)
by: Park, Chanwoo, et al.
Published: (2024)
Similar Items
-
Reward-Robust RLHF in LLMs
by: Yan, Yuzi, et al.
Published: (2024) -
3D-Properties: Identifying Challenges in DPO and Charting a Path Forward
by: Yan, Yuzi, et al.
Published: (2024) -
Exploring the LLM Journey from Cognition to Expression with Linear Representations
by: Yan, Yuzi, et al.
Published: (2024) -
Policy Filtration for RLHF to Mitigate Noise in Reward Models
by: Zhang, Chuheng, et al.
Published: (2024) -
RLHF Workflow: From Reward Modeling to Online RLHF
by: Dong, Hanze, et al.
Published: (2024)