Confidence as a Reward: Transforming LLMs into Reward Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Du, He, Li, Bowen, Xie, Chengxing, Gao, Chang, Chen, Kai, Tao, Dacheng |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Intra-Trajectory Consistency for Reward Modeling
par: Zhou, Chaoyang, et autres
Publié: (2025)
par: Zhou, Chaoyang, et autres
Publié: (2025)
InfoRM: Mitigating Reward Hacking in RLHF via Information-Theoretic Reward Modeling
par: Miao, Yuchun, et autres
Publié: (2024)
par: Miao, Yuchun, et autres
Publié: (2024)
Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs
par: Liu, Chris Yuhao, et autres
Publié: (2024)
par: Liu, Chris Yuhao, et autres
Publié: (2024)
Why Self-Rewarding Works: Theoretical Guarantees for Iterative Alignment of Language Models
par: Fu, Shi, et autres
Publié: (2026)
par: Fu, Shi, et autres
Publié: (2026)
CAMEL: Confidence-Gated Reflection for Reward Modeling
par: Zhu, Zirui, et autres
Publié: (2026)
par: Zhu, Zirui, et autres
Publié: (2026)
Teaching LLMs to Abstain via Fine-Grained Semantic Confidence Reward
par: An, Hao, et autres
Publié: (2025)
par: An, Hao, et autres
Publié: (2025)
PACR: Progressively Ascending Confidence Reward for LLM Reasoning
par: Yoon, Eunseop, et autres
Publié: (2025)
par: Yoon, Eunseop, et autres
Publié: (2025)
Beyond Correctness: Confidence-Aware Reward Modeling for Enhancing Large Language Model Reasoning
par: He, Qianxi, et autres
Publié: (2025)
par: He, Qianxi, et autres
Publié: (2025)
Joint Reward Modeling: Internalizing Chain-of-Thought for Efficient Visual Reward Models
par: Yang, Yankai, et autres
Publié: (2026)
par: Yang, Yankai, et autres
Publié: (2026)
Structural Reward Model: Enhancing Interpretability, Efficiency, and Scalability in Reward Modeling
par: Liu, Xiaoyu, et autres
Publié: (2025)
par: Liu, Xiaoyu, et autres
Publié: (2025)
On the Transformations across Reward Model, Parameter Update, and In-Context Prompt
par: Cai, Deng, et autres
Publié: (2024)
par: Cai, Deng, et autres
Publié: (2024)
Reward-Robust RLHF in LLMs
par: Yan, Yuzi, et autres
Publié: (2024)
par: Yan, Yuzi, et autres
Publié: (2024)
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models
par: Zhou, Xin, et autres
Publié: (2025)
par: Zhou, Xin, et autres
Publié: (2025)
Text2Reward: Reward Shaping with Language Models for Reinforcement Learning
par: Xie, Tianbao, et autres
Publié: (2023)
par: Xie, Tianbao, et autres
Publié: (2023)
Confidence-aware Reward Optimization for Fine-tuning Text-to-Image Models
par: Kim, Kyuyoung, et autres
Publié: (2024)
par: Kim, Kyuyoung, et autres
Publié: (2024)
Learning in Context, Guided by Choice: A Reward-Free Paradigm for Reinforcement Learning with Transformers
par: Dong, Juncheng, et autres
Publié: (2026)
par: Dong, Juncheng, et autres
Publié: (2026)
Reward Model Perspectives: Whose Opinions Do Reward Models Reward?
par: Elle
Publié: (2025)
par: Elle
Publié: (2025)
GRAM: A Generative Foundation Reward Model for Reward Generalization
par: Wang, Chenglong, et autres
Publié: (2025)
par: Wang, Chenglong, et autres
Publié: (2025)
When Right Meets Wrong: Bilateral Context Conditioning with Reward-Confidence Correction for GRPO
par: Li, Yu, et autres
Publié: (2026)
par: Li, Yu, et autres
Publié: (2026)
DocReward: A Document Reward Model for Structuring and Stylizing
par: Liu, Junpeng, et autres
Publié: (2025)
par: Liu, Junpeng, et autres
Publié: (2025)
Towards Understanding the Influence of Reward Margin on Preference Model Performance
par: Qin, Bowen, et autres
Publié: (2024)
par: Qin, Bowen, et autres
Publié: (2024)
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution
par: Li, Jiahui, et autres
Publié: (2024)
par: Li, Jiahui, et autres
Publié: (2024)
MemReward: Graph-Based Experience Memory for LLM Reward Prediction with Limited Labels
par: Luo, Tianyang, et autres
Publié: (2026)
par: Luo, Tianyang, et autres
Publié: (2026)
OS-Themis: A Scalable Critic Framework for Generalist GUI Rewards
par: Li, Zehao, et autres
Publié: (2026)
par: Li, Zehao, et autres
Publié: (2026)
Reward-Augmented Data Enhances Direct Preference Alignment of LLMs
par: Zhang, Shenao, et autres
Publié: (2024)
par: Zhang, Shenao, et autres
Publié: (2024)
RLHF in an SFT Way: From Optimal Solution to Reward-Weighted Alignment
par: Du, Yuhao, et autres
Publié: (2025)
par: Du, Yuhao, et autres
Publié: (2025)
CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward
par: Liu, Shudong, et autres
Publié: (2025)
par: Liu, Shudong, et autres
Publié: (2025)
Reward Bound for Behavioral Guarantee of Model-based Planning Agents
par: An, Zhiyu, et autres
Publié: (2024)
par: An, Zhiyu, et autres
Publié: (2024)
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
par: Wang, Chaoqi, et autres
Publié: (2025)
par: Wang, Chaoqi, et autres
Publié: (2025)
LoRe: Personalizing LLMs via Low-Rank Reward Modeling
par: Bose, Avinandan, et autres
Publié: (2025)
par: Bose, Avinandan, et autres
Publié: (2025)
A Human-Like Reasoning Framework for Multi-Phases Planning Task with Large Language Models
par: Xie, Chengxing, et autres
Publié: (2024)
par: Xie, Chengxing, et autres
Publié: (2024)
CMI-RewardBench: Evaluating Music Reward Models with Compositional Multimodal Instruction
par: Ma, Yinghao, et autres
Publié: (2026)
par: Ma, Yinghao, et autres
Publié: (2026)
Direct Advantage Regression: Aligning LLMs with Online AI Reward
par: He, Li, et autres
Publié: (2025)
par: He, Li, et autres
Publié: (2025)
Self-Evolved Reward Learning for LLMs
par: Huang, Chenghua, et autres
Publié: (2024)
par: Huang, Chenghua, et autres
Publié: (2024)
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning
par: Zhang, Kongcheng, et autres
Publié: (2025)
par: Zhang, Kongcheng, et autres
Publié: (2025)
SemiReward: A General Reward Model for Semi-supervised Learning
par: Li, Siyuan, et autres
Publié: (2023)
par: Li, Siyuan, et autres
Publié: (2023)
SAFER: Probing Safety in Reward Models with Sparse Autoencoder
par: Shi, Wei, et autres
Publié: (2025)
par: Shi, Wei, et autres
Publié: (2025)
GUI-PRA: Process Reward Agent for GUI Tasks
par: Xiong, Tao, et autres
Publié: (2025)
par: Xiong, Tao, et autres
Publié: (2025)
Decoupling Reasoning and Confidence: Resurrecting Calibration in Reinforcement Learning from Verifiable Rewards
par: Ma, Zhengzhao, et autres
Publié: (2026)
par: Ma, Zhengzhao, et autres
Publié: (2026)
WildReward: Learning Reward Models from In-the-Wild Human Interactions
par: Peng, Hao, et autres
Publié: (2026)
par: Peng, Hao, et autres
Publié: (2026)
Documents similaires
-
Intra-Trajectory Consistency for Reward Modeling
par: Zhou, Chaoyang, et autres
Publié: (2025) -
InfoRM: Mitigating Reward Hacking in RLHF via Information-Theoretic Reward Modeling
par: Miao, Yuchun, et autres
Publié: (2024) -
Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs
par: Liu, Chris Yuhao, et autres
Publié: (2024) -
Why Self-Rewarding Works: Theoretical Guarantees for Iterative Alignment of Language Models
par: Fu, Shi, et autres
Publié: (2026) -
CAMEL: Confidence-Gated Reflection for Reward Modeling
par: Zhu, Zirui, et autres
Publié: (2026)