Process Reward Models That Think
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Khalifa, Muhammad, Agarwal, Rishabh, Logeswaran, Lajanugen, Kim, Jaekyeom, Peng, Hao, Lee, Moontae, Lee, Honglak, Wang, Lu |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Auto-Intent: Automated Intent Discovery and Self-Exploration for Large Language Model Web Agents
par: Kim, Jaekyeom, et autres
Publié: (2024)
par: Kim, Jaekyeom, et autres
Publié: (2024)
GRACE: Discriminator-Guided Chain-of-Thought Reasoning
par: Khalifa, Muhammad, et autres
Publié: (2023)
par: Khalifa, Muhammad, et autres
Publié: (2023)
Gaming the Judge: Unfaithful Chain-of-Thought Can Undermine Agent Evaluation
par: Khalifa, Muhammad, et autres
Publié: (2026)
par: Khalifa, Muhammad, et autres
Publié: (2026)
Small Language Models Need Strong Verifiers to Self-Correct Reasoning
par: Zhang, Yunxiang, et autres
Publié: (2024)
par: Zhang, Yunxiang, et autres
Publié: (2024)
SPRIG: Improving Large Language Model Performance by System Prompt Optimization
par: Zhang, Lechen, et autres
Publié: (2024)
par: Zhang, Lechen, et autres
Publié: (2024)
When "A Helpful Assistant" Is Not Really Helpful: Personas in System Prompts Do Not Improve Performances of Large Language Models
par: Zheng, Mingqian, et autres
Publié: (2023)
par: Zheng, Mingqian, et autres
Publié: (2023)
AutoGuide: Automated Generation and Selection of Context-Aware Guidelines for Large Language Model Agents
par: Fu, Yao, et autres
Publié: (2024)
par: Fu, Yao, et autres
Publié: (2024)
Cross-Lingual Prompt Steerability: Towards Accurate and Robust LLM Behavior across Languages
par: Zhang, Lechen, et autres
Publié: (2025)
par: Zhang, Lechen, et autres
Publié: (2025)
Visual Test-time Scaling for GUI Agent Grounding
par: Luo, Tiange, et autres
Publié: (2025)
par: Luo, Tiange, et autres
Publié: (2025)
Selective LoRA for Visual Tokens and Attention Heads
par: Luo, Tiange, et autres
Publié: (2025)
par: Luo, Tiange, et autres
Publié: (2025)
MLRC-Bench: Can Language Agents Solve Machine Learning Research Challenges?
par: Zhang, Yunxiang, et autres
Publié: (2025)
par: Zhang, Yunxiang, et autres
Publié: (2025)
Scaling Web Agent Training through Automatic Data Generation and Fine-grained Evaluation
par: Logeswaran, Lajanugen, et autres
Publié: (2026)
par: Logeswaran, Lajanugen, et autres
Publié: (2026)
Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents
par: Jang, Yunseok, et autres
Publié: (2025)
par: Jang, Yunseok, et autres
Publié: (2025)
Countdown-Code: A Testbed for Studying The Emergence and Generalization of Reward Hacking in RLVR
par: Khalifa, Muhammad, et autres
Publié: (2026)
par: Khalifa, Muhammad, et autres
Publié: (2026)
Understanding the Capabilities and Limitations of Large Language Models for Cultural Commonsense
par: Shen, Siqi, et autres
Publié: (2024)
par: Shen, Siqi, et autres
Publié: (2024)
Revisiting LLM Value Probing Strategies: Are They Robust and Expressive?
par: Shen, Siqi, et autres
Publié: (2025)
par: Shen, Siqi, et autres
Publié: (2025)
You don't need a personality test to know these models are unreliable: Assessing the Reliability of Large Language Models on Psychometric Instruments
par: Shu, Bangzhao, et autres
Publié: (2023)
par: Shu, Bangzhao, et autres
Publié: (2023)
Efficient Process Reward Modeling via Contrastive Mutual Information
par: Lee, Nakyung, et autres
Publié: (2026)
par: Lee, Nakyung, et autres
Publié: (2026)
Source-Aware Training Enables Knowledge Attribution in Language Models
par: Khalifa, Muhammad, et autres
Publié: (2024)
par: Khalifa, Muhammad, et autres
Publié: (2024)
Mitigating Biases for Instruction-following Language Models via Bias Neurons Elimination
par: Yang, Nakyeong, et autres
Publié: (2023)
par: Yang, Nakyeong, et autres
Publié: (2023)
Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization
par: Kim, Sunghwan, et autres
Publié: (2025)
par: Kim, Sunghwan, et autres
Publié: (2025)
From Faithfulness to Correctness: Generative Reward Models that Think Critically
par: Ma, Qiyao, et autres
Publié: (2025)
par: Ma, Qiyao, et autres
Publié: (2025)
On the Robustness of Reward Models for Language Model Alignment
par: Hong, Jiwoo, et autres
Publié: (2025)
par: Hong, Jiwoo, et autres
Publié: (2025)
More Bang for the Buck: Process Reward Modeling with Entropy-Driven Uncertainty
par: Cao, Lang, et autres
Publié: (2025)
par: Cao, Lang, et autres
Publié: (2025)
Evaluating Robustness of Reward Models for Mathematical Reasoning
par: Kim, Sunghwan, et autres
Publié: (2024)
par: Kim, Sunghwan, et autres
Publié: (2024)
M-RewardBench: Evaluating Reward Models in Multilingual Settings
par: Gureja, Srishti, et autres
Publié: (2024)
par: Gureja, Srishti, et autres
Publié: (2024)
Process Reinforcement through Implicit Rewards
par: Cui, Ganqu, et autres
Publié: (2025)
par: Cui, Ganqu, et autres
Publié: (2025)
Unlocking Multimodal Mathematical Reasoning via Process Reward Model
par: Luo, Ruilin, et autres
Publié: (2025)
par: Luo, Ruilin, et autres
Publié: (2025)
SafeDPO: A Simple Approach to Direct Preference Optimization with Enhanced Safety
par: Kim, Geon-Hyeong, et autres
Publié: (2025)
par: Kim, Geon-Hyeong, et autres
Publié: (2025)
Interactive and Expressive Code-Augmented Planning with Large Language Models
par: Liu, Anthony Z., et autres
Publié: (2024)
par: Liu, Anthony Z., et autres
Publié: (2024)
Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models
par: Noukhovitch, Michael, et autres
Publié: (2024)
par: Noukhovitch, Michael, et autres
Publié: (2024)
The Lessons of Developing Process Reward Models in Mathematical Reasoning
par: Zhang, Zhenru, et autres
Publié: (2025)
par: Zhang, Zhenru, et autres
Publié: (2025)
Rewarding Graph Reasoning Process makes LLMs more Generalized Reasoners
par: Peng, Miao, et autres
Publié: (2025)
par: Peng, Miao, et autres
Publié: (2025)
Think Inside the JSON: Reinforcement Strategy for Strict LLM Schema Adherence
par: Agarwal, Bhavik, et autres
Publié: (2025)
par: Agarwal, Bhavik, et autres
Publié: (2025)
What Makes a Reward Model a Good Teacher? An Optimization Perspective
par: Razin, Noam, et autres
Publié: (2025)
par: Razin, Noam, et autres
Publié: (2025)
Process Rewards with Learned Reliability
par: Li, Jinyuan, et autres
Publié: (2026)
par: Li, Jinyuan, et autres
Publié: (2026)
BRiTE: Bootstrapping Reinforced Thinking Process to Enhance Language Model Reasoning
par: Zhong, Han, et autres
Publié: (2025)
par: Zhong, Han, et autres
Publié: (2025)
Think Before You Prune: Self-Reflective Structured Pruning for Reasoning Language Models
par: Wang, Ziyan, et autres
Publié: (2025)
par: Wang, Ziyan, et autres
Publié: (2025)
SPARK: Stepwise Process-Aware Rewards for Reference-Free Reinforcement Learning
par: Rahman, Salman, et autres
Publié: (2025)
par: Rahman, Salman, et autres
Publié: (2025)
Self-Guided Process Reward Optimization with Redefined Step-wise Advantage for Process Reinforcement Learning
par: Fei, Wu, et autres
Publié: (2025)
par: Fei, Wu, et autres
Publié: (2025)
Documents similaires
-
Auto-Intent: Automated Intent Discovery and Self-Exploration for Large Language Model Web Agents
par: Kim, Jaekyeom, et autres
Publié: (2024) -
GRACE: Discriminator-Guided Chain-of-Thought Reasoning
par: Khalifa, Muhammad, et autres
Publié: (2023) -
Gaming the Judge: Unfaithful Chain-of-Thought Can Undermine Agent Evaluation
par: Khalifa, Muhammad, et autres
Publié: (2026) -
Small Language Models Need Strong Verifiers to Self-Correct Reasoning
par: Zhang, Yunxiang, et autres
Publié: (2024) -
SPRIG: Improving Large Language Model Performance by System Prompt Optimization
par: Zhang, Lechen, et autres
Publié: (2024)