Do We Need to Verify Step by Step? Rethinking Process Supervision from a Theoretical Perspective
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jia, Zeyu, Rakhlin, Alexander, Xie, Tengyang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Outcome-Based Online Reinforcement Learning: Algorithms and Fundamental Limits
von: Chen, Fan, et al.
Veröffentlicht: (2025)
von: Chen, Fan, et al.
Veröffentlicht: (2025)
Trajectory Bellman Residual Minimization: A Simple Value-Based Method for LLM Reasoning
von: Yuan, Yurun, et al.
Veröffentlicht: (2025)
von: Yuan, Yurun, et al.
Veröffentlicht: (2025)
Offline Reinforcement Learning: Role of State Aggregation and Trajectory Data
von: Jia, Zeyu, et al.
Veröffentlicht: (2024)
von: Jia, Zeyu, et al.
Veröffentlicht: (2024)
Do We Need Frontier Models to Verify Mathematical Proofs?
von: Naik, Aaditya, et al.
Veröffentlicht: (2026)
von: Naik, Aaditya, et al.
Veröffentlicht: (2026)
Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF
von: Xie, Tengyang, et al.
Veröffentlicht: (2024)
von: Xie, Tengyang, et al.
Veröffentlicht: (2024)
Do Latent-CoT Models Think Step-by-Step? A Mechanistic Study on Sequential Reasoning Tasks
von: Liang, Jia, et al.
Veröffentlicht: (2026)
von: Liang, Jia, et al.
Veröffentlicht: (2026)
Learning to Generate Formally Verifiable Step-by-Step Logic Reasoning via Structured Formal Intermediaries
von: Chen, Luoxin, et al.
Veröffentlicht: (2026)
von: Chen, Luoxin, et al.
Veröffentlicht: (2026)
Offline Reinforcement Learning in Large State Spaces: Algorithms and Guarantees
von: Jiang, Nan, et al.
Veröffentlicht: (2025)
von: Jiang, Nan, et al.
Veröffentlicht: (2025)
SSPO: Self-traced Step-wise Preference Optimization for Process Supervision and Reasoning Compression
von: Xu, Yuyang, et al.
Veröffentlicht: (2025)
von: Xu, Yuyang, et al.
Veröffentlicht: (2025)
Multi-Step Likelihood-Ratio Correction for Reinforcement Learning with Verifiable Rewards
von: Yoon, Deokgyu, et al.
Veröffentlicht: (2026)
von: Yoon, Deokgyu, et al.
Veröffentlicht: (2026)
Reinforce LLM Reasoning through Multi-Agent Reflection
von: Yuan, Yurun, et al.
Veröffentlicht: (2025)
von: Yuan, Yurun, et al.
Veröffentlicht: (2025)
Breaking the Capability Ceiling of LLM Post-Training by Reintroducing Markov States
von: Yuan, Yurun, et al.
Veröffentlicht: (2026)
von: Yuan, Yurun, et al.
Veröffentlicht: (2026)
LLM Reasoning with Process Rewards for Outcome-Guided Steps
von: Rezaei, Mohammad, et al.
Veröffentlicht: (2026)
von: Rezaei, Mohammad, et al.
Veröffentlicht: (2026)
The Power of Resets in Online Reinforcement Learning
von: Mhammedi, Zakaria, et al.
Veröffentlicht: (2024)
von: Mhammedi, Zakaria, et al.
Veröffentlicht: (2024)
The Signal is in the Steps: Local Scoring for Reasoning Data Selection
von: Just, Hoang Anh, et al.
Veröffentlicht: (2025)
von: Just, Hoang Anh, et al.
Veröffentlicht: (2025)
Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
von: Wang, Peiyi, et al.
Veröffentlicht: (2023)
von: Wang, Peiyi, et al.
Veröffentlicht: (2023)
POLCA: Stochastic Generative Optimization with LLM
von: Ren, Xuanfei, et al.
Veröffentlicht: (2026)
von: Ren, Xuanfei, et al.
Veröffentlicht: (2026)
Boosting Deductive Reasoning with Step Signals In RLHF
von: Li, Jialian, et al.
Veröffentlicht: (2024)
von: Li, Jialian, et al.
Veröffentlicht: (2024)
Watch Every Step! LLM Agent Learning via Iterative Step-Level Process Refinement
von: Xiong, Weimin, et al.
Veröffentlicht: (2024)
von: Xiong, Weimin, et al.
Veröffentlicht: (2024)
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
von: Lai, Xin, et al.
Veröffentlicht: (2024)
von: Lai, Xin, et al.
Veröffentlicht: (2024)
EngTrace: A Symbolic Benchmark for Verifiable Process Supervision of Engineering Reasoning
von: Gull, Ayesha, et al.
Veröffentlicht: (2025)
von: Gull, Ayesha, et al.
Veröffentlicht: (2025)
Do We Need Adam? Surprisingly Strong and Sparse Reinforcement Learning with SGD in LLMs
von: Mukherjee, Sagnik, et al.
Veröffentlicht: (2026)
von: Mukherjee, Sagnik, et al.
Veröffentlicht: (2026)
Decision Making in Changing Environments: Robustness, Query-Based Learning, and Differential Privacy
von: Chen, Fan, et al.
Veröffentlicht: (2025)
von: Chen, Fan, et al.
Veröffentlicht: (2025)
Step-wise Adaptive Integration of Supervised Fine-tuning and Reinforcement Learning for Task-Specific LLMs
von: Chen, Jack, et al.
Veröffentlicht: (2025)
von: Chen, Jack, et al.
Veröffentlicht: (2025)
One Step Forward and K Steps Back: Better Reasoning with Denoising Recursion Models
von: Cameron, Chris, et al.
Veröffentlicht: (2026)
von: Cameron, Chris, et al.
Veröffentlicht: (2026)
Pencil Puzzle Bench: A Benchmark for Multi-Step Verifiable Reasoning
von: Waugh, Justin
Veröffentlicht: (2026)
von: Waugh, Justin
Veröffentlicht: (2026)
DreamPRM-Code: Function-as-Step Process Reward Model with Label Correction for LLM Coding
von: Zhang, Ruiyi, et al.
Veröffentlicht: (2025)
von: Zhang, Ruiyi, et al.
Veröffentlicht: (2025)
Can We Verify Step by Step for Incorrect Answer Detection?
von: Xu, Xin, et al.
Veröffentlicht: (2024)
von: Xu, Xin, et al.
Veröffentlicht: (2024)
Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier Math
von: Pandit, Shrey, et al.
Veröffentlicht: (2025)
von: Pandit, Shrey, et al.
Veröffentlicht: (2025)
Supervised Reinforcement Learning: From Expert Trajectories to Step-wise Reasoning
von: Deng, Yihe, et al.
Veröffentlicht: (2025)
von: Deng, Yihe, et al.
Veröffentlicht: (2025)
Effect of a Process Mining based Pre-processing Step in Prediction of the Critical Health Outcomes
von: Ashrafi, Negin, et al.
Veröffentlicht: (2024)
von: Ashrafi, Negin, et al.
Veröffentlicht: (2024)
Diffusion World Model: Future Modeling Beyond Step-by-Step Rollout for Offline Reinforcement Learning
von: Ding, Zihan, et al.
Veröffentlicht: (2024)
von: Ding, Zihan, et al.
Veröffentlicht: (2024)
GRPO-VPS: Enhancing Group Relative Policy Optimization with Verifiable Process Supervision for Effective Reasoning
von: Wang, Jingyi, et al.
Veröffentlicht: (2026)
von: Wang, Jingyi, et al.
Veröffentlicht: (2026)
Towards a Theoretical Understanding of Synthetic Data in LLM Post-Training: A Reverse-Bottleneck Perspective
von: Gan, Zeyu, et al.
Veröffentlicht: (2024)
von: Gan, Zeyu, et al.
Veröffentlicht: (2024)
Step-by-Step Causality: Transparent Causal Discovery with Multi-Agent Tree-Query and Adversarial Confidence Estimation
von: Ding, Ziyi, et al.
Veröffentlicht: (2026)
von: Ding, Ziyi, et al.
Veröffentlicht: (2026)
The Hidden Signal of Verifier Strictness: Controlling and Improving Step-Wise Verification via Selective Latent Steering
von: Zhou, Yefan, et al.
Veröffentlicht: (2026)
von: Zhou, Yefan, et al.
Veröffentlicht: (2026)
Were RNNs All We Needed?
von: Feng, Leo, et al.
Veröffentlicht: (2024)
von: Feng, Leo, et al.
Veröffentlicht: (2024)
Generative AI Models for Different Steps in Architectural Design: A Literature Review
von: Li, Chengyuan, et al.
Veröffentlicht: (2024)
von: Li, Chengyuan, et al.
Veröffentlicht: (2024)
SBSC: Step-By-Step Coding for Improving Mathematical Olympiad Performance
von: Singh, Kunal, et al.
Veröffentlicht: (2025)
von: Singh, Kunal, et al.
Veröffentlicht: (2025)
Step-size Optimization for Continual Learning
von: Degris, Thomas, et al.
Veröffentlicht: (2024)
von: Degris, Thomas, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Outcome-Based Online Reinforcement Learning: Algorithms and Fundamental Limits
von: Chen, Fan, et al.
Veröffentlicht: (2025) -
Trajectory Bellman Residual Minimization: A Simple Value-Based Method for LLM Reasoning
von: Yuan, Yurun, et al.
Veröffentlicht: (2025) -
Offline Reinforcement Learning: Role of State Aggregation and Trajectory Data
von: Jia, Zeyu, et al.
Veröffentlicht: (2024) -
Do We Need Frontier Models to Verify Mathematical Proofs?
von: Naik, Aaditya, et al.
Veröffentlicht: (2026) -
Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF
von: Xie, Tengyang, et al.
Veröffentlicht: (2024)