DreamPRM-Code: Function-as-Step Process Reward Model with Label Correction for LLM Coding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Ruiyi, Qin, Peijia, Cao, Qi, Xie, Pengtao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DreamPRM: Domain-Reweighted Process Reward Model for Multimodal Reasoning
von: Cao, Qi, et al.
Veröffentlicht: (2025)
von: Cao, Qi, et al.
Veröffentlicht: (2025)
FunPRM: Function-as-Step Process Reward Model with Meta Reward Correction for Code Generation
von: Zhang, Ruiyi, et al.
Veröffentlicht: (2026)
von: Zhang, Ruiyi, et al.
Veröffentlicht: (2026)
DreamPRM-1.5: Unlocking the Potential of Each Instance for Multimodal Process Reward Model Training
von: Cao, Qi, et al.
Veröffentlicht: (2025)
von: Cao, Qi, et al.
Veröffentlicht: (2025)
DAJ: Data-Reweighted LLM Judge for Test-Time Scaling in Code Generation
von: Qin, Peijia, et al.
Veröffentlicht: (2026)
von: Qin, Peijia, et al.
Veröffentlicht: (2026)
BiDoRA: Bi-level Optimization-Based Weight-Decomposed Low-Rank Adaptation
von: Qin, Peijia, et al.
Veröffentlicht: (2024)
von: Qin, Peijia, et al.
Veröffentlicht: (2024)
AutoLoRA: Automatically Tuning Matrix Ranks in Low-Rank Adaptation Based on Meta Learning
von: Zhang, Ruiyi, et al.
Veröffentlicht: (2024)
von: Zhang, Ruiyi, et al.
Veröffentlicht: (2024)
AIBuildAI: An AI Agent for Automatically Building AI Models
von: Zhang, Ruiyi, et al.
Veröffentlicht: (2026)
von: Zhang, Ruiyi, et al.
Veröffentlicht: (2026)
Multi-Turn Code Generation Through Single-Step Rewards
von: Jain, Arnav Kumar, et al.
Veröffentlicht: (2025)
von: Jain, Arnav Kumar, et al.
Veröffentlicht: (2025)
ReCode: Reinforcing Code Generation with Reasoning-Process Rewards
von: Fan, Lishui, et al.
Veröffentlicht: (2025)
von: Fan, Lishui, et al.
Veröffentlicht: (2025)
Functional Entropy: Predicting Functional Correctness in LLM-Generated Code with Uncertainty Quantification
von: Bouchard, Dylan, et al.
Veröffentlicht: (2026)
von: Bouchard, Dylan, et al.
Veröffentlicht: (2026)
Dreaming in Code for Curriculum Learning in Open-Ended Worlds
von: Mitsides, Konstantinos, et al.
Veröffentlicht: (2026)
von: Mitsides, Konstantinos, et al.
Veröffentlicht: (2026)
AgentPRM: Process Reward Models for LLM Agents via Step-Wise Promise and Progress
von: Xi, Zhiheng, et al.
Veröffentlicht: (2025)
von: Xi, Zhiheng, et al.
Veröffentlicht: (2025)
Models Under SCOPE: Scalable and Controllable Routing via Pre-hoc Reasoning
von: Cao, Qi, et al.
Veröffentlicht: (2026)
von: Cao, Qi, et al.
Veröffentlicht: (2026)
SBSC: Step-By-Step Coding for Improving Mathematical Olympiad Performance
von: Singh, Kunal, et al.
Veröffentlicht: (2025)
von: Singh, Kunal, et al.
Veröffentlicht: (2025)
Reasoning Through Execution: Unifying Process and Outcome Rewards for Code Generation
von: Yu, Zhuohao, et al.
Veröffentlicht: (2024)
von: Yu, Zhuohao, et al.
Veröffentlicht: (2024)
ATLAS: Agentic Test-time Learning-to-Allocate Scaling
von: Qin, Peijia, et al.
Veröffentlicht: (2026)
von: Qin, Peijia, et al.
Veröffentlicht: (2026)
Self-Guided Process Reward Optimization with Redefined Step-wise Advantage for Process Reinforcement Learning
von: Fei, Wu, et al.
Veröffentlicht: (2025)
von: Fei, Wu, et al.
Veröffentlicht: (2025)
LLM Reasoning as Trajectories: Step-Specific Representation Geometry and Correctness Signals
von: Sun, Lihao, et al.
Veröffentlicht: (2026)
von: Sun, Lihao, et al.
Veröffentlicht: (2026)
Process Reward Models That Think
von: Khalifa, Muhammad, et al.
Veröffentlicht: (2025)
von: Khalifa, Muhammad, et al.
Veröffentlicht: (2025)
BiLoRA: A Bi-level Optimization Framework for Overfitting-Resilient Low-Rank Adaptation of Large Pre-trained Models
von: Qiang, Rushi, et al.
Veröffentlicht: (2024)
von: Qiang, Rushi, et al.
Veröffentlicht: (2024)
CodeScaler: Scaling Code LLM Training and Test-Time Inference via Reward Models
von: Zhu, Xiao, et al.
Veröffentlicht: (2026)
von: Zhu, Xiao, et al.
Veröffentlicht: (2026)
Automated Rewards via LLM-Generated Progress Functions
von: Sarukkai, Vishnu, et al.
Veröffentlicht: (2024)
von: Sarukkai, Vishnu, et al.
Veröffentlicht: (2024)
CodeRefine: A Pipeline for Enhancing LLM-Generated Code Implementations of Research Papers
von: Trofimova, Ekaterina, et al.
Veröffentlicht: (2024)
von: Trofimova, Ekaterina, et al.
Veröffentlicht: (2024)
Inverse Reinforcement Learning with Dynamic Reward Scaling for LLM Alignment
von: Cheng, Ruoxi, et al.
Veröffentlicht: (2025)
von: Cheng, Ruoxi, et al.
Veröffentlicht: (2025)
Let the Code LLM Edit Itself When You Edit the Code
von: He, Zhenyu, et al.
Veröffentlicht: (2024)
von: He, Zhenyu, et al.
Veröffentlicht: (2024)
Countdown-Code: A Testbed for Studying The Emergence and Generalization of Reward Hacking in RLVR
von: Khalifa, Muhammad, et al.
Veröffentlicht: (2026)
von: Khalifa, Muhammad, et al.
Veröffentlicht: (2026)
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment
von: Zhang, Jiazheng, et al.
Veröffentlicht: (2025)
von: Zhang, Jiazheng, et al.
Veröffentlicht: (2025)
Generalizable and Stable Finetuning of Pretrained Language Models on Low-Resource Texts
von: Somayajula, Sai Ashish, et al.
Veröffentlicht: (2024)
von: Somayajula, Sai Ashish, et al.
Veröffentlicht: (2024)
The Lessons of Developing Process Reward Models in Mathematical Reasoning
von: Zhang, Zhenru, et al.
Veröffentlicht: (2025)
von: Zhang, Zhenru, et al.
Veröffentlicht: (2025)
Watch Every Step! LLM Agent Learning via Iterative Step-Level Process Refinement
von: Xiong, Weimin, et al.
Veröffentlicht: (2024)
von: Xiong, Weimin, et al.
Veröffentlicht: (2024)
On Mitigating Code LLM Hallucinations with API Documentation
von: Jain, Nihal, et al.
Veröffentlicht: (2024)
von: Jain, Nihal, et al.
Veröffentlicht: (2024)
Code Comprehension then Auditing for Unsupervised LLM Evaluation
von: Patel, Bhrij, et al.
Veröffentlicht: (2024)
von: Patel, Bhrij, et al.
Veröffentlicht: (2024)
LLM Reasoning with Process Rewards for Outcome-Guided Steps
von: Rezaei, Mohammad, et al.
Veröffentlicht: (2026)
von: Rezaei, Mohammad, et al.
Veröffentlicht: (2026)
CodeARC: Benchmarking Reasoning Capabilities of LLM Agents for Inductive Program Synthesis
von: Wei, Anjiang, et al.
Veröffentlicht: (2025)
von: Wei, Anjiang, et al.
Veröffentlicht: (2025)
Your Reward Function for RL is Your Best PRM for Search: Unifying RL and Search-Based TTS
von: Jin, Can, et al.
Veröffentlicht: (2025)
von: Jin, Can, et al.
Veröffentlicht: (2025)
Planning In Natural Language Improves LLM Search For Code Generation
von: Wang, Evan, et al.
Veröffentlicht: (2024)
von: Wang, Evan, et al.
Veröffentlicht: (2024)
Verbal Process Supervision Elicits Better Coding Agents
von: Chen, Hao-Yuan, et al.
Veröffentlicht: (2025)
von: Chen, Hao-Yuan, et al.
Veröffentlicht: (2025)
VeriCoder: Enhancing LLM-Based RTL Code Generation through Functional Correctness Validation
von: Wei, Anjiang, et al.
Veröffentlicht: (2025)
von: Wei, Anjiang, et al.
Veröffentlicht: (2025)
More Bang for the Buck: Process Reward Modeling with Entropy-Driven Uncertainty
von: Cao, Lang, et al.
Veröffentlicht: (2025)
von: Cao, Lang, et al.
Veröffentlicht: (2025)
Process Rewards with Learned Reliability
von: Li, Jinyuan, et al.
Veröffentlicht: (2026)
von: Li, Jinyuan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
DreamPRM: Domain-Reweighted Process Reward Model for Multimodal Reasoning
von: Cao, Qi, et al.
Veröffentlicht: (2025) -
FunPRM: Function-as-Step Process Reward Model with Meta Reward Correction for Code Generation
von: Zhang, Ruiyi, et al.
Veröffentlicht: (2026) -
DreamPRM-1.5: Unlocking the Potential of Each Instance for Multimodal Process Reward Model Training
von: Cao, Qi, et al.
Veröffentlicht: (2025) -
DAJ: Data-Reweighted LLM Judge for Test-Time Scaling in Code Generation
von: Qin, Peijia, et al.
Veröffentlicht: (2026) -
BiDoRA: Bi-level Optimization-Based Weight-Decomposed Low-Rank Adaptation
von: Qin, Peijia, et al.
Veröffentlicht: (2024)