Reward Models Can Improve Themselves: Reward-Guided Adversarial Failure Mode Discovery for Robust Reward Modeling
Fuente:
arXiv
Saved in:
| Main Authors: | Pathmanathan, Pankayaraj, Huang, Furong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time Alignment
by: Xu, Yuancheng, et al.
Published: (2024)
by: Xu, Yuancheng, et al.
Published: (2024)
Deliberative Alignment is Deep, but Uncertainty Remains: Inference time safety improvement in reasoning via attribution of unsafe behavior to base model
by: Pathmanathan, Pankayaraj, et al.
Published: (2026)
by: Pathmanathan, Pankayaraj, et al.
Published: (2026)
Is poisoning a real threat to LLM alignment? Maybe more so than you think
by: Pathmanathan, Pankayaraj, et al.
Published: (2024)
by: Pathmanathan, Pankayaraj, et al.
Published: (2024)
RRM: Robust Reward Model Training Mitigates Reward Hacking
by: Liu, Tianqi, et al.
Published: (2024)
by: Liu, Tianqi, et al.
Published: (2024)
Reward Model Perspectives: Whose Opinions Do Reward Models Reward?
by: Elle
Published: (2025)
by: Elle
Published: (2025)
Improving Reward Models with Synthetic Critiques
by: Ye, Zihuiwen, et al.
Published: (2024)
by: Ye, Zihuiwen, et al.
Published: (2024)
Reward Reasoning Model
by: Guo, Jiaxin, et al.
Published: (2025)
by: Guo, Jiaxin, et al.
Published: (2025)
RewardBench 2: Advancing Reward Model Evaluation
by: Malik, Saumya, et al.
Published: (2025)
by: Malik, Saumya, et al.
Published: (2025)
Long-form RewardBench: Evaluating Reward Models for Long-form Generation
by: Huang, Hui, et al.
Published: (2026)
by: Huang, Hui, et al.
Published: (2026)
Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs
by: Liu, Chris Yuhao, et al.
Published: (2024)
by: Liu, Chris Yuhao, et al.
Published: (2024)
EditReward: A Human-Aligned Reward Model for Instruction-Guided Image Editing
by: Wu, Keming, et al.
Published: (2025)
by: Wu, Keming, et al.
Published: (2025)
DocReward: A Document Reward Model for Structuring and Stylizing
by: Liu, Junpeng, et al.
Published: (2025)
by: Liu, Junpeng, et al.
Published: (2025)
reWordBench: Benchmarking and Improving the Robustness of Reward Models with Transformed Inputs
by: Wu, Zhaofeng, et al.
Published: (2025)
by: Wu, Zhaofeng, et al.
Published: (2025)
Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision
by: Pala, Tej Deep, et al.
Published: (2025)
by: Pala, Tej Deep, et al.
Published: (2025)
On the Robustness of Reward Models for Language Model Alignment
by: Hong, Jiwoo, et al.
Published: (2025)
by: Hong, Jiwoo, et al.
Published: (2025)
Libra: Assessing and Improving Reward Model by Learning to Think
by: Zhou, Meng, et al.
Published: (2025)
by: Zhou, Meng, et al.
Published: (2025)
Dynamic and Generalizable Process Reward Modeling
by: Yin, Zhangyue, et al.
Published: (2025)
by: Yin, Zhangyue, et al.
Published: (2025)
GRAM: A Generative Foundation Reward Model for Reward Generalization
by: Wang, Chenglong, et al.
Published: (2025)
by: Wang, Chenglong, et al.
Published: (2025)
RewardAnything: Generalizable Principle-Following Reward Models
by: Yu, Zhuohao, et al.
Published: (2025)
by: Yu, Zhuohao, et al.
Published: (2025)
Tiny Reward Models
by: Pan, Sarah
Published: (2025)
by: Pan, Sarah
Published: (2025)
Reward Auditor: Inference on Reward Modeling Suitability in Real-World Perturbed Scenarios
by: Zang, Jianxiang, et al.
Published: (2025)
by: Zang, Jianxiang, et al.
Published: (2025)
ENCORE: Entropy-guided Reward Composition for Multi-head Safety Reward Models
by: Li, Xiaomin, et al.
Published: (2025)
by: Li, Xiaomin, et al.
Published: (2025)
The Bidirectional Process Reward Model
by: Zhang, Lingyin, et al.
Published: (2025)
by: Zhang, Lingyin, et al.
Published: (2025)
Tool-Augmented Reward Modeling
by: Li, Lei, et al.
Published: (2023)
by: Li, Lei, et al.
Published: (2023)
Fine-tuning Language Models with Generative Adversarial Reward Modelling
by: Yu, Zhang Ze, et al.
Published: (2023)
by: Yu, Zhang Ze, et al.
Published: (2023)
Rubric-Guided Process Reward for Stepwise Model Routing
by: Ye, Shenghao, et al.
Published: (2026)
by: Ye, Shenghao, et al.
Published: (2026)
WildReward: Learning Reward Models from In-the-Wild Human Interactions
by: Peng, Hao, et al.
Published: (2026)
by: Peng, Hao, et al.
Published: (2026)
Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization
by: Ma, Qiyao, et al.
Published: (2026)
by: Ma, Qiyao, et al.
Published: (2026)
Evaluating Robustness of Reward Models for Mathematical Reasoning
by: Kim, Sunghwan, et al.
Published: (2024)
by: Kim, Sunghwan, et al.
Published: (2024)
LLM-as-a-Judge & Reward Model: What They Can and Cannot Do
by: Son, Guijin, et al.
Published: (2024)
by: Son, Guijin, et al.
Published: (2024)
Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards
by: Wu, Xiaobao
Published: (2025)
by: Wu, Xiaobao
Published: (2025)
Reward-Augmented Decoding: Efficient Controlled Text Generation With a Unidirectional Reward Model
by: Deng, Haikang, et al.
Published: (2023)
by: Deng, Haikang, et al.
Published: (2023)
Reuse Your Rewards: Reward Model Transfer for Zero-Shot Cross-Lingual Alignment
by: Wu, Zhaofeng, et al.
Published: (2024)
by: Wu, Zhaofeng, et al.
Published: (2024)
Energy-Based Reward Models for Robust Language Model Alignment
by: Lochab, Anamika, et al.
Published: (2025)
by: Lochab, Anamika, et al.
Published: (2025)
Robust Preference Optimization through Reward Model Distillation
by: Fisch, Adam, et al.
Published: (2024)
by: Fisch, Adam, et al.
Published: (2024)
Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization
by: Kim, Sunghwan, et al.
Published: (2025)
by: Kim, Sunghwan, et al.
Published: (2025)
M-RewardBench: Evaluating Reward Models in Multilingual Settings
by: Gureja, Srishti, et al.
Published: (2024)
by: Gureja, Srishti, et al.
Published: (2024)
R-PRM: Reasoning-Driven Process Reward Modeling
by: She, Shuaijie, et al.
Published: (2025)
by: She, Shuaijie, et al.
Published: (2025)
Universal Adversarial Suffixes for Language Models Using Reinforcement Learning with Calibrated Reward
by: Soor, Sampriti, et al.
Published: (2025)
by: Soor, Sampriti, et al.
Published: (2025)
Self-Rewarding Language Models
by: Yuan, Weizhe, et al.
Published: (2024)
by: Yuan, Weizhe, et al.
Published: (2024)
Similar Items
-
GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time Alignment
by: Xu, Yuancheng, et al.
Published: (2024) -
Deliberative Alignment is Deep, but Uncertainty Remains: Inference time safety improvement in reasoning via attribution of unsafe behavior to base model
by: Pathmanathan, Pankayaraj, et al.
Published: (2026) -
Is poisoning a real threat to LLM alignment? Maybe more so than you think
by: Pathmanathan, Pankayaraj, et al.
Published: (2024) -
RRM: Robust Reward Model Training Mitigates Reward Hacking
by: Liu, Tianqi, et al.
Published: (2024) -
Reward Model Perspectives: Whose Opinions Do Reward Models Reward?
by: Elle
Published: (2025)