Process Reward Models for LLM Agents: Practical Framework and Directions
Fuente:
arXiv
Saved in:
| Main Author: | Choudhury, Sanjiban |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Better than Your Teacher: LLM Agents that learn from Privileged AI Feedback
by: Choudhury, Sanjiban, et al.
Published: (2024)
by: Choudhury, Sanjiban, et al.
Published: (2024)
Aligning LLMs with Domain Invariant Reward Models
by: Wu, David, et al.
Published: (2025)
by: Wu, David, et al.
Published: (2025)
Efficient Imitation under Misspecification
by: Espinosa-Dice, Nicolas, et al.
Published: (2025)
by: Espinosa-Dice, Nicolas, et al.
Published: (2025)
Multi-Turn Code Generation Through Single-Step Rewards
by: Jain, Arnav Kumar, et al.
Published: (2025)
by: Jain, Arnav Kumar, et al.
Published: (2025)
InteRACT: Transformer Models for Human Intent Prediction Conditioned on Robot Actions
by: Kedia, Kushal, et al.
Published: (2023)
by: Kedia, Kushal, et al.
Published: (2023)
Motion Tracks: A Unified Representation for Human-Robot Transfer in Few-Shot Imitation Learning
by: Ren, Juntao, et al.
Published: (2025)
by: Ren, Juntao, et al.
Published: (2025)
Hybrid Inverse Reinforcement Learning
by: Ren, Juntao, et al.
Published: (2024)
by: Ren, Juntao, et al.
Published: (2024)
One-Shot Imitation under Mismatched Execution
by: Kedia, Kushal, et al.
Published: (2024)
by: Kedia, Kushal, et al.
Published: (2024)
Imitation Learning via Focused Satisficing
by: Shah, Rushit N., et al.
Published: (2025)
by: Shah, Rushit N., et al.
Published: (2025)
Non-Adversarial Inverse Reinforcement Learning via Successor Feature Matching
by: Jain, Arnav Kumar, et al.
Published: (2024)
by: Jain, Arnav Kumar, et al.
Published: (2024)
LLM Reasoning with Process Rewards for Outcome-Guided Steps
by: Rezaei, Mohammad, et al.
Published: (2026)
by: Rezaei, Mohammad, et al.
Published: (2026)
Adversarial Training for Process Reward Models
by: Juneja, Gurusha, et al.
Published: (2025)
by: Juneja, Gurusha, et al.
Published: (2025)
Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use
by: Thaman, Kunvar
Published: (2026)
by: Thaman, Kunvar
Published: (2026)
GRPO is Secretly a Process Reward Model
by: Sullivan, Michael, et al.
Published: (2025)
by: Sullivan, Michael, et al.
Published: (2025)
Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes
by: Kobalczyk, Katarzyna, et al.
Published: (2024)
by: Kobalczyk, Katarzyna, et al.
Published: (2024)
Distributional Process Reward Models: Calibrated Prediction of Future Rewards via Conditional Optimal Transport
by: Ma, Rachel, et al.
Published: (2026)
by: Ma, Rachel, et al.
Published: (2026)
X-Sim: Cross-Embodiment Learning via Real-to-Sim-to-Real
by: Dan, Prithwish, et al.
Published: (2025)
by: Dan, Prithwish, et al.
Published: (2025)
Process Reward Models That Think
by: Khalifa, Muhammad, et al.
Published: (2025)
by: Khalifa, Muhammad, et al.
Published: (2025)
SignalLLM: A General-Purpose LLM Agent Framework for Automated Signal Processing
by: Ke, Junlong, et al.
Published: (2025)
by: Ke, Junlong, et al.
Published: (2025)
Efficient Process Reward Model Training via Active Learning
by: Duan, Keyu, et al.
Published: (2025)
by: Duan, Keyu, et al.
Published: (2025)
Optimal Transport for LLM Reward Modeling from Noisy Preference
by: Pan, Licheng, et al.
Published: (2026)
by: Pan, Licheng, et al.
Published: (2026)
PROF: An LLM-based Reward Code Preference Optimization Framework for Offline Imitation Learning
by: Sun, Shengjie, et al.
Published: (2025)
by: Sun, Shengjie, et al.
Published: (2025)
Beyond Numeric Rewards: In-Context Dueling Bandits with LLM Agents
by: Xia, Fanzeng, et al.
Published: (2024)
by: Xia, Fanzeng, et al.
Published: (2024)
DreamPRM: Domain-Reweighted Process Reward Model for Multimodal Reasoning
by: Cao, Qi, et al.
Published: (2025)
by: Cao, Qi, et al.
Published: (2025)
DreamPRM-Code: Function-as-Step Process Reward Model with Label Correction for LLM Coding
by: Zhang, Ruiyi, et al.
Published: (2025)
by: Zhang, Ruiyi, et al.
Published: (2025)
Accelerating LLM Reasoning via Early Rejection with Partial Reward Modeling
by: Cheshmi, Seyyed Saeid, et al.
Published: (2025)
by: Cheshmi, Seyyed Saeid, et al.
Published: (2025)
AgentRM: Enhancing Agent Generalization with Reward Modeling
by: Xia, Yu, et al.
Published: (2025)
by: Xia, Yu, et al.
Published: (2025)
Same Signal, Opposite Meaning: Direction-Informed Adaptive Learning for LLM Agents
by: Li, Ziming, et al.
Published: (2026)
by: Li, Ziming, et al.
Published: (2026)
Know What You Don't Know: Uncertainty Calibration of Process Reward Models
by: Park, Young-Jin, et al.
Published: (2025)
by: Park, Young-Jin, et al.
Published: (2025)
MemReward: Graph-Based Experience Memory for LLM Reward Prediction with Limited Labels
by: Luo, Tianyang, et al.
Published: (2026)
by: Luo, Tianyang, et al.
Published: (2026)
Reward-Augmented Data Enhances Direct Preference Alignment of LLMs
by: Zhang, Shenao, et al.
Published: (2024)
by: Zhang, Shenao, et al.
Published: (2024)
ProgAgent:A Continual RL Agent with Progress-Aware Rewards
by: Tan, Jinzhou, et al.
Published: (2026)
by: Tan, Jinzhou, et al.
Published: (2026)
RewardUQ: A Unified Framework for Uncertainty-Aware Reward Models
by: Yang, Daniel, et al.
Published: (2026)
by: Yang, Daniel, et al.
Published: (2026)
Model-Based Reinforcement Learning in Discrete-Action Non-Markovian Reward Decision Processes
by: Trapasso, Alessandro, et al.
Published: (2025)
by: Trapasso, Alessandro, et al.
Published: (2025)
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning
by: Li, Mengqi, et al.
Published: (2025)
by: Li, Mengqi, et al.
Published: (2025)
Rethinking Rubric Generation for Improving LLM Judge and Reward Modeling for Open-ended Tasks
by: Shen, William F., et al.
Published: (2026)
by: Shen, William F., et al.
Published: (2026)
Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
by: Rafailov, Rafael, et al.
Published: (2024)
by: Rafailov, Rafael, et al.
Published: (2024)
$ξ$-DPO: Direct Preference Optimization via Ratio Reward Margin
by: Fan, Zhengyuan, et al.
Published: (2026)
by: Fan, Zhengyuan, et al.
Published: (2026)
Robotouille: An Asynchronous Planning Benchmark for LLM Agents
by: Gonzalez-Pumariega, Gonzalo, et al.
Published: (2025)
by: Gonzalez-Pumariega, Gonzalo, et al.
Published: (2025)
Debate as Reward: A Multi-Agent Reward System for Scientific Ideation via RL Post-Training
by: Salimi, Moein, et al.
Published: (2026)
by: Salimi, Moein, et al.
Published: (2026)
Similar Items
-
Better than Your Teacher: LLM Agents that learn from Privileged AI Feedback
by: Choudhury, Sanjiban, et al.
Published: (2024) -
Aligning LLMs with Domain Invariant Reward Models
by: Wu, David, et al.
Published: (2025) -
Efficient Imitation under Misspecification
by: Espinosa-Dice, Nicolas, et al.
Published: (2025) -
Multi-Turn Code Generation Through Single-Step Rewards
by: Jain, Arnav Kumar, et al.
Published: (2025) -
InteRACT: Transformer Models for Human Intent Prediction Conditioned on Robot Actions
by: Kedia, Kushal, et al.
Published: (2023)