Is PRM Necessary? Problem-Solving RL Implicitly Induces PRM Capability in LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Feng, Zhangying, Chen, Qianglong, Lu, Ning, Li, Yongqian, Cheng, Siqi, Peng, Shuangmu, Tang, Duyu, Liu, Shengcai, Zhang, Zhirui |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
TIM-PRM: Verifying multimodal reasoning with Tool-Integrated PRM
di: Kuang, Peng, et al.
Pubblicazione: (2025)
di: Kuang, Peng, et al.
Pubblicazione: (2025)
PRM-BAS: Enhancing Multimodal Reasoning through PRM-guided Beam Annealing Search
di: Hu, Pengfei, et al.
Pubblicazione: (2025)
di: Hu, Pengfei, et al.
Pubblicazione: (2025)
rePIRL: Learn PRM with Inverse RL for LLM Reasoning
di: Wu, Xian, et al.
Pubblicazione: (2026)
di: Wu, Xian, et al.
Pubblicazione: (2026)
Balneo and PRM Research Journal
Pubblicazione: (2021)
Pubblicazione: (2021)
Your Reward Function for RL is Your Best PRM for Search: Unifying RL and Search-Based TTS
di: Jin, Can, et al.
Pubblicazione: (2025)
di: Jin, Can, et al.
Pubblicazione: (2025)
Limits of PRM-Guided Tree Search for Mathematical Reasoning with LLMs
di: Cinquin, Tristan, et al.
Pubblicazione: (2025)
di: Cinquin, Tristan, et al.
Pubblicazione: (2025)
PRM: Photometric Stereo based Large Reconstruction Model
di: Ge, Wenhang, et al.
Pubblicazione: (2024)
di: Ge, Wenhang, et al.
Pubblicazione: (2024)
R-PRM: Reasoning-Driven Process Reward Modeling
di: She, Shuaijie, et al.
Pubblicazione: (2025)
di: She, Shuaijie, et al.
Pubblicazione: (2025)
Recursive Structure of Hulls of PRM Codes
di: Song, Yufeng, et al.
Pubblicazione: (2026)
di: Song, Yufeng, et al.
Pubblicazione: (2026)
FreePRM: Training Process Reward Models Without Ground Truth Process Labels
di: Sun, Lin, et al.
Pubblicazione: (2025)
di: Sun, Lin, et al.
Pubblicazione: (2025)
VisualPRM: An Effective Process Reward Model for Multimodal Reasoning
di: Wang, Weiyun, et al.
Pubblicazione: (2025)
di: Wang, Weiyun, et al.
Pubblicazione: (2025)
Next Step Mobile: Strategy, Services, & PRM
di: Thomas, Lisa Carlucci
Pubblicazione: (2012)
di: Thomas, Lisa Carlucci
Pubblicazione: (2012)
Cutting corners in muscle measurements with ISarcoPRM!
di: Ahmad J. Abdulsalam
Pubblicazione: (2024)
di: Ahmad J. Abdulsalam
Pubblicazione: (2024)
Optimal Aggregation of LLM and PRM Signals for Efficient Test-Time Scaling
di: Kuang, Peng, et al.
Pubblicazione: (2025)
di: Kuang, Peng, et al.
Pubblicazione: (2025)
DW-A-PRM: A Dynamic Weighted Planner
di: Wang, Siyuan, et al.
Pubblicazione: (2025)
di: Wang, Siyuan, et al.
Pubblicazione: (2025)
ReasonFlux-PRM: Trajectory-Aware PRMs for Long Chain-of-Thought Reasoning in LLMs
di: Zou, Jiaru, et al.
Pubblicazione: (2025)
di: Zou, Jiaru, et al.
Pubblicazione: (2025)
ContextPRM: Leveraging Contextual Coherence for multi-domain Test-Time Scaling
di: Zhang, Haotian, et al.
Pubblicazione: (2025)
di: Zhang, Haotian, et al.
Pubblicazione: (2025)
Med-PRM: Medical Reasoning Models with Stepwise, Guideline-verified Process Rewards
di: Yun, Jaehoon, et al.
Pubblicazione: (2025)
di: Yun, Jaehoon, et al.
Pubblicazione: (2025)
ToolPRM: Fine-Grained Inference Scaling of Structured Outputs for Function Calling
di: Lin, Jianghao, et al.
Pubblicazione: (2025)
di: Lin, Jianghao, et al.
Pubblicazione: (2025)
Efficient PRM Training Data Synthesis via Formal Verification
di: Kamoi, Ryo, et al.
Pubblicazione: (2025)
di: Kamoi, Ryo, et al.
Pubblicazione: (2025)
Less is More: Understanding Word-level Textual Adversarial Attack via n-gram Frequency Descend
di: Lu, Ning, et al.
Pubblicazione: (2023)
di: Lu, Ning, et al.
Pubblicazione: (2023)
MM-PRM: Enhancing Multimodal Mathematical Reasoning with Scalable Step-Level Supervision
di: Du, Lingxiao, et al.
Pubblicazione: (2025)
di: Du, Lingxiao, et al.
Pubblicazione: (2025)
SecCodePRM: A Process Reward Model for Code Security
di: Yu, Weichen, et al.
Pubblicazione: (2026)
di: Yu, Weichen, et al.
Pubblicazione: (2026)
DreamPRM: Domain-Reweighted Process Reward Model for Multimodal Reasoning
di: Cao, Qi, et al.
Pubblicazione: (2025)
di: Cao, Qi, et al.
Pubblicazione: (2025)
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets
di: Lu, Ning, et al.
Pubblicazione: (2025)
di: Lu, Ning, et al.
Pubblicazione: (2025)
PRM-as-a-Judge: A Dense Evaluation Paradigm for Fine-Grained Robotic Auditing
di: Ji, Yuheng, et al.
Pubblicazione: (2026)
di: Ji, Yuheng, et al.
Pubblicazione: (2026)
TaTToo: Tool-Grounded Thinking PRM for Test-Time Scaling in Tabular Reasoning
di: Zou, Jiaru, et al.
Pubblicazione: (2025)
di: Zou, Jiaru, et al.
Pubblicazione: (2025)
SwarmPRM: Probabilistic Roadmap Motion Planning for Large-Scale Swarm Robotic Systems
di: Hu, Yunze, et al.
Pubblicazione: (2024)
di: Hu, Yunze, et al.
Pubblicazione: (2024)
Fin-PRM: A Domain-Specialized Process Reward Model for Financial Reasoning in Large Language Models
di: Zhu, Jie, et al.
Pubblicazione: (2025)
di: Zhu, Jie, et al.
Pubblicazione: (2025)
PRM path smoothening by circular arc fillet method for mobile robot navigation
di: Ouach, Meral Kılıçarslan, et al.
Pubblicazione: (2021)
di: Ouach, Meral Kılıçarslan, et al.
Pubblicazione: (2021)
PRM ‐Star: An Enhanced Probabilistic Roadmap Algorithm With Adaptive Sampling and Path Optimization
di: Oluwaseun O. Martins, et al.
Pubblicazione: (2026)
di: Oluwaseun O. Martins, et al.
Pubblicazione: (2026)
Evaluación de híbridos experimentales de maíz del PRM en Centroamérica
di: Mario Fuentes
Pubblicazione: (2003)
di: Mario Fuentes
Pubblicazione: (2003)
AutoPRM: Automating Procedural Supervision for Multi-Step Reasoning via Controllable Question Decomposition
di: Chen, Zhaorun, et al.
Pubblicazione: (2024)
di: Chen, Zhaorun, et al.
Pubblicazione: (2024)
H-PRM: A Pluggable Hotword Pre-Retrieval Module for Various Speech Recognition Systems
di: Dai, Huangyu, et al.
Pubblicazione: (2025)
di: Dai, Huangyu, et al.
Pubblicazione: (2025)
GroundedPRM: Tree-Guided and Fidelity-Aware Process Reward Modeling for Step-Level Reasoning
di: Zhang, Yao, et al.
Pubblicazione: (2025)
di: Zhang, Yao, et al.
Pubblicazione: (2025)
FG-PRM: Fine-grained Hallucination Detection and Mitigation in Language Model Mathematical Reasoning
di: Li, Ruosen, et al.
Pubblicazione: (2024)
di: Li, Ruosen, et al.
Pubblicazione: (2024)
Accurate and Diverse LLM Mathematical Reasoning via Automated PRM-Guided GFlowNets
di: Younsi, Adam, et al.
Pubblicazione: (2025)
di: Younsi, Adam, et al.
Pubblicazione: (2025)
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training
di: Du, Pengfei
Pubblicazione: (2025)
di: Du, Pengfei
Pubblicazione: (2025)
GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning
di: Zhang, Jianghangfan, et al.
Pubblicazione: (2025)
di: Zhang, Jianghangfan, et al.
Pubblicazione: (2025)
VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data
di: Zeng, Thomas, et al.
Pubblicazione: (2025)
di: Zeng, Thomas, et al.
Pubblicazione: (2025)
Documenti analoghi
-
TIM-PRM: Verifying multimodal reasoning with Tool-Integrated PRM
di: Kuang, Peng, et al.
Pubblicazione: (2025) -
PRM-BAS: Enhancing Multimodal Reasoning through PRM-guided Beam Annealing Search
di: Hu, Pengfei, et al.
Pubblicazione: (2025) -
rePIRL: Learn PRM with Inverse RL for LLM Reasoning
di: Wu, Xian, et al.
Pubblicazione: (2026) -
Balneo and PRM Research Journal
Pubblicazione: (2021) -
Your Reward Function for RL is Your Best PRM for Search: Unifying RL and Search-Based TTS
di: Jin, Can, et al.
Pubblicazione: (2025)