Uncertainty-Aware Step-wise Verification with Generative Reward Models
Fuente:
arXiv
Saved in:
| Main Authors: | Ye, Zihuiwen, Melo, Luckeciano Carvalho, Kaddar, Younesse, Blunsom, Phil, Staton, Sam, Gal, Yarin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Likelihood hacking in probabilistic program synthesis
by: Karwowski, Jacek, et al.
Published: (2026)
by: Karwowski, Jacek, et al.
Published: (2026)
A model of stochastic memoization and name generation in probabilistic programming: categorical semantics via monads on presheaf categories
by: Kaddar, Younesse, et al.
Published: (2023)
by: Kaddar, Younesse, et al.
Published: (2023)
Improving Reward Models with Synthetic Critiques
by: Ye, Zihuiwen, et al.
Published: (2024)
by: Ye, Zihuiwen, et al.
Published: (2024)
Deep Bayesian Active Learning for Preference Modeling in Large Language Models
by: Melo, Luckeciano C., et al.
Published: (2024)
by: Melo, Luckeciano C., et al.
Published: (2024)
Uncertainty Quantification for LLM Function-Calling
by: Ye, Zihuiwen, et al.
Published: (2026)
by: Ye, Zihuiwen, et al.
Published: (2026)
Probabilistic programming interfaces for random graphs: Markov categories, graphons, and nominal sets
by: Ackerman, Nathanael L., et al.
Published: (2023)
by: Ackerman, Nathanael L., et al.
Published: (2023)
Iterative Deployment Improves Planning Skills in LLMs
by: Corrêa, Augusto B., et al.
Published: (2025)
by: Corrêa, Augusto B., et al.
Published: (2025)
Human Feedback is not Gold Standard
by: Hosking, Tom, et al.
Published: (2023)
by: Hosking, Tom, et al.
Published: (2023)
Stabilizing Policy Gradients for Sample-Efficient Reinforcement Learning in LLM Reasoning
by: Melo, Luckeciano C., et al.
Published: (2025)
by: Melo, Luckeciano C., et al.
Published: (2025)
Temporal-Difference Variational Continual Learning
by: Melo, Luckeciano C., et al.
Published: (2024)
by: Melo, Luckeciano C., et al.
Published: (2024)
Amortizing intractable inference in large language models
by: Hu, Edward J., et al.
Published: (2023)
by: Hu, Edward J., et al.
Published: (2023)
Language Models Change Facts Based on the Way You Talk
by: Kearney, Matthew, et al.
Published: (2025)
by: Kearney, Matthew, et al.
Published: (2025)
Kernel Language Entropy: Fine-grained Uncertainty Quantification for LLMs from Semantic Similarities
by: Nikitin, Alexander, et al.
Published: (2024)
by: Nikitin, Alexander, et al.
Published: (2024)
Full-Step-DPO: Self-Supervised Preference Optimization with Step-wise Rewards for Mathematical Reasoning
by: Xu, Huimin, et al.
Published: (2025)
by: Xu, Huimin, et al.
Published: (2025)
SPADER: Step-wise Peer Advantage with Diversity-Aware Exploration Rewards for Multi-Answer Question Answering
by: Shi, Qiming, et al.
Published: (2026)
by: Shi, Qiming, et al.
Published: (2026)
Higher Order Automatic Differentiation of Higher Order Functions
by: Huot, Mathieu, et al.
Published: (2021)
by: Huot, Mathieu, et al.
Published: (2021)
Do Multilingual LLMs Think In English?
by: Schut, Lisa, et al.
Published: (2025)
by: Schut, Lisa, et al.
Published: (2025)
Compositional imprecise probability
by: Liell-Cock, Jack, et al.
Published: (2024)
by: Liell-Cock, Jack, et al.
Published: (2024)
In-Context Learning Learns Label Relationships but Is Not Conventional Learning
by: Kossen, Jannik, et al.
Published: (2023)
by: Kossen, Jannik, et al.
Published: (2023)
Stop Rewarding Hallucinated Steps: Faithfulness-Aware Step-Level Reinforcement Learning for Small Reasoning Models
by: Nie, Shuo, et al.
Published: (2026)
by: Nie, Shuo, et al.
Published: (2026)
From Tokens to Steps: Verification-Aware Speculative Decoding for Efficient Multi-Step Reasoning
by: Purohit, Kiran, et al.
Published: (2026)
by: Purohit, Kiran, et al.
Published: (2026)
Rope to Nope and Back Again: A New Hybrid Attention Strategy
by: Yang, Bowen, et al.
Published: (2025)
by: Yang, Bowen, et al.
Published: (2025)
RewardUQ: A Unified Framework for Uncertainty-Aware Reward Models
by: Yang, Daniel, et al.
Published: (2026)
by: Yang, Daniel, et al.
Published: (2026)
Process Reward Models Meet Planning: Generating Precise and Scalable Datasets for Step-Level Rewards
by: Pisano, Raffaele, et al.
Published: (2026)
by: Pisano, Raffaele, et al.
Published: (2026)
Fine-Tuning Large Language Models to Appropriately Abstain with Semantic Entropy
by: Tjandra, Benedict Aaron, et al.
Published: (2024)
by: Tjandra, Benedict Aaron, et al.
Published: (2024)
The Curse of Recursion: Training on Generated Data Makes Models Forget
by: Shumailov, Ilia, et al.
Published: (2023)
by: Shumailov, Ilia, et al.
Published: (2023)
Self-Guided Process Reward Optimization with Redefined Step-wise Advantage for Process Reinforcement Learning
by: Fei, Wu, et al.
Published: (2025)
by: Fei, Wu, et al.
Published: (2025)
Bayesian Preference Elicitation with Language Models
by: Handa, Kunal, et al.
Published: (2024)
by: Handa, Kunal, et al.
Published: (2024)
Evaluating Step-by-Step Reasoning through Symbolic Verification
by: Zhang, Yi-Fan, et al.
Published: (2022)
by: Zhang, Yi-Fan, et al.
Published: (2022)
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification
by: Liu, Rui, et al.
Published: (2026)
by: Liu, Rui, et al.
Published: (2026)
CARMO: Dynamic Criteria Generation for Context-Aware Reward Modelling
by: Gupta, Taneesh, et al.
Published: (2024)
by: Gupta, Taneesh, et al.
Published: (2024)
Can Large Language Models Faithfully Express Their Intrinsic Uncertainty in Words?
by: Yona, Gal, et al.
Published: (2024)
by: Yona, Gal, et al.
Published: (2024)
SOD: Step-wise On-policy Distillation for Small Language Model Agents
by: Zhong, Qiyong, et al.
Published: (2026)
by: Zhong, Qiyong, et al.
Published: (2026)
Reason Only When Needed: Efficient Generative Reward Modeling via Model-Internal Uncertainty
by: Xue, Chao, et al.
Published: (2026)
by: Xue, Chao, et al.
Published: (2026)
InfoQuest: Evaluating Multi-Turn Dialogue Agents for Open-Ended Conversations with Hidden Context
by: de Oliveira, Bryan L. M., et al.
Published: (2025)
by: de Oliveira, Bryan L. M., et al.
Published: (2025)
Evaluating & Reducing Deceptive Dialogue From Language Models with Multi-turn RL
by: Abdulhai, Marwa, et al.
Published: (2025)
by: Abdulhai, Marwa, et al.
Published: (2025)
Uncertainty-Aware Variational Reward Factorization via Probabilistic Preference Bases for LLM Personalization
by: Lee, Gyuseok, et al.
Published: (2026)
by: Lee, Gyuseok, et al.
Published: (2026)
STEPER: Step-wise Knowledge Distillation for Enhancing Reasoning Ability in Multi-Step Retrieval-Augmented Language Models
by: Lee, Kyumin, et al.
Published: (2025)
by: Lee, Kyumin, et al.
Published: (2025)
Towards Hierarchical Multi-Step Reward Models for Enhanced Reasoning in Large Language Models
by: Wang, Teng, et al.
Published: (2025)
by: Wang, Teng, et al.
Published: (2025)
MUSIC: MUlti-Step Instruction Contrast for Multi-Turn Reward Models
by: Li, Wenzhe, et al.
Published: (2025)
by: Li, Wenzhe, et al.
Published: (2025)
Similar Items
-
Likelihood hacking in probabilistic program synthesis
by: Karwowski, Jacek, et al.
Published: (2026) -
A model of stochastic memoization and name generation in probabilistic programming: categorical semantics via monads on presheaf categories
by: Kaddar, Younesse, et al.
Published: (2023) -
Improving Reward Models with Synthetic Critiques
by: Ye, Zihuiwen, et al.
Published: (2024) -
Deep Bayesian Active Learning for Preference Modeling in Large Language Models
by: Melo, Luckeciano C., et al.
Published: (2024) -
Uncertainty Quantification for LLM Function-Calling
by: Ye, Zihuiwen, et al.
Published: (2026)