RMB: Comprehensively Benchmarking Reward Models in LLM Alignment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Enyu, Zheng, Guodong, Wang, Binghai, Xi, Zhiheng, Dou, Shihan, Bao, Rong, Shen, Wei, Xiong, Limao, Fan, Jessica, Mou, Yurong, Zheng, Rui, Gui, Tao, Zhang, Qi, Huang, Xuanjing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MetaRM: Shifted Distributions Alignment via Meta-Learning
von: Dou, Shihan, et al.
Veröffentlicht: (2024)
von: Dou, Shihan, et al.
Veröffentlicht: (2024)
StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback
von: Dou, Shihan, et al.
Veröffentlicht: (2024)
von: Dou, Shihan, et al.
Veröffentlicht: (2024)
Improving RL Exploration for LLM Reasoning through Retrospective Replay
von: Dou, Shihan, et al.
Veröffentlicht: (2025)
von: Dou, Shihan, et al.
Veröffentlicht: (2025)
Steering LLMs via Scalable Interactive Oversight
von: Zhou, Enyu, et al.
Veröffentlicht: (2026)
von: Zhou, Enyu, et al.
Veröffentlicht: (2026)
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment
von: Zhang, Jiazheng, et al.
Veröffentlicht: (2025)
von: Zhang, Jiazheng, et al.
Veröffentlicht: (2025)
From Scores to Preferences: Redefining MOS Benchmarking for Speech Quality Reward Modeling
von: Cao, Yifei, et al.
Veröffentlicht: (2025)
von: Cao, Yifei, et al.
Veröffentlicht: (2025)
JFTA-Bench: Evaluate LLM's Ability of Tracking and Analyzing Malfunctions Using Fault Trees
von: Wang, Yuhui, et al.
Veröffentlicht: (2026)
von: Wang, Yuhui, et al.
Veröffentlicht: (2026)
Subspace Defense: Discarding Adversarial Perturbations by Learning a Subspace for Clean Signals
von: Zheng, Rui, et al.
Veröffentlicht: (2024)
von: Zheng, Rui, et al.
Veröffentlicht: (2024)
MM-Doc-R1: Training Agents for Long Document Visual Question Answering through Multi-turn Reinforcement Learning
von: Lin, Jiahang, et al.
Veröffentlicht: (2026)
von: Lin, Jiahang, et al.
Veröffentlicht: (2026)
Enhancing LLM-based Search Agents via Contribution Weighted Group Relative Policy Optimization
von: Wang, Junzhe, et al.
Veröffentlicht: (2026)
von: Wang, Junzhe, et al.
Veröffentlicht: (2026)
SafeAligner: Safety Alignment against Jailbreak Attacks via Response Disparity Guidance
von: Huang, Caishuang, et al.
Veröffentlicht: (2024)
von: Huang, Caishuang, et al.
Veröffentlicht: (2024)
Aligning Large Language Models from Self-Reference AI Feedback with one General Principle
von: Bao, Rong, et al.
Veröffentlicht: (2024)
von: Bao, Rong, et al.
Veröffentlicht: (2024)
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training
von: Pan, Chengjun, et al.
Veröffentlicht: (2026)
von: Pan, Chengjun, et al.
Veröffentlicht: (2026)
Toward Optimal LLM Alignments Using Two-Player Games
von: Zheng, Rui, et al.
Veröffentlicht: (2024)
von: Zheng, Rui, et al.
Veröffentlicht: (2024)
Secrets of RLHF in Large Language Models Part II: Reward Modeling
von: Wang, Binghai, et al.
Veröffentlicht: (2024)
von: Wang, Binghai, et al.
Veröffentlicht: (2024)
AgentPRM: Process Reward Models for LLM Agents via Step-Wise Promise and Progress
von: Xi, Zhiheng, et al.
Veröffentlicht: (2025)
von: Xi, Zhiheng, et al.
Veröffentlicht: (2025)
LoRAMoE: Alleviate World Knowledge Forgetting in Large Language Models via MoE-Style Plugin
von: Dou, Shihan, et al.
Veröffentlicht: (2023)
von: Dou, Shihan, et al.
Veröffentlicht: (2023)
SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model
von: Zhang, Yongting, et al.
Veröffentlicht: (2024)
von: Zhang, Yongting, et al.
Veröffentlicht: (2024)
Better Process Supervision with Bi-directional Rewarding Signals
von: Chen, Wenxiang, et al.
Veröffentlicht: (2025)
von: Chen, Wenxiang, et al.
Veröffentlicht: (2025)
Self-Polish: Enhance Reasoning in Large Language Models via Problem Refinement
von: Xi, Zhiheng, et al.
Veröffentlicht: (2023)
von: Xi, Zhiheng, et al.
Veröffentlicht: (2023)
What's Wrong with Your Code Generated by Large Language Models? An Extensive Study
von: Dou, Shihan, et al.
Veröffentlicht: (2024)
von: Dou, Shihan, et al.
Veröffentlicht: (2024)
Pre-Trained Policy Discriminators are General Reward Models
von: Dou, Shihan, et al.
Veröffentlicht: (2025)
von: Dou, Shihan, et al.
Veröffentlicht: (2025)
EasyJailbreak: A Unified Framework for Jailbreaking Large Language Models
von: Zhou, Weikang, et al.
Veröffentlicht: (2024)
von: Zhou, Weikang, et al.
Veröffentlicht: (2024)
Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses
von: Lin, Jiahang, et al.
Veröffentlicht: (2026)
von: Lin, Jiahang, et al.
Veröffentlicht: (2026)
ChartE$^{3}$: A Comprehensive Benchmark for End-to-End Chart Editing
von: Li, Shuo, et al.
Veröffentlicht: (2026)
von: Li, Shuo, et al.
Veröffentlicht: (2026)
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents
von: Shen, Yujiong, et al.
Veröffentlicht: (2026)
von: Shen, Yujiong, et al.
Veröffentlicht: (2026)
LLMEval-Med: A Real-world Clinical Benchmark for Medical LLMs with Physician Validation
von: Zhang, Ming, et al.
Veröffentlicht: (2025)
von: Zhang, Ming, et al.
Veröffentlicht: (2025)
LLMEval-Logic: A Solver-Verified Chinese Benchmark for Logical Reasoning of LLMs with Adversarial Hardening
von: Zhang, Ming, et al.
Veröffentlicht: (2026)
von: Zhang, Ming, et al.
Veröffentlicht: (2026)
Unlocking the Essence of Beauty: Advanced Aesthetic Reasoning with Relative-Absolute Policy Optimization
von: Liu, Boyang, et al.
Veröffentlicht: (2025)
von: Liu, Boyang, et al.
Veröffentlicht: (2025)
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models
von: Zhou, Xin, et al.
Veröffentlicht: (2025)
von: Zhou, Xin, et al.
Veröffentlicht: (2025)
RoCoIns: Enhancing Robustness of Large Language Models through Code-Style Instructions
von: Zhang, Yuansen, et al.
Veröffentlicht: (2024)
von: Zhang, Yuansen, et al.
Veröffentlicht: (2024)
VRPO: Rethinking Value Modeling for Robust RL Training under Noisy Supervision
von: Zhu, Dingwei, et al.
Veröffentlicht: (2025)
von: Zhu, Dingwei, et al.
Veröffentlicht: (2025)
SpeechRole: A Large-Scale Dataset and Benchmark for Evaluating Speech Role-Playing Agents
von: Jiang, Changhao, et al.
Veröffentlicht: (2025)
von: Jiang, Changhao, et al.
Veröffentlicht: (2025)
Mitigating Object Hallucinations in MLLMs via Multi-Frequency Perturbations
von: Li, Shuo, et al.
Veröffentlicht: (2025)
von: Li, Shuo, et al.
Veröffentlicht: (2025)
Advancing Translation Preference Modeling with RLHF: A Step Towards Cost-Effective Solution
von: Xu, Nuo, et al.
Veröffentlicht: (2024)
von: Xu, Nuo, et al.
Veröffentlicht: (2024)
Mitigating Tail Narrowing in LLM Self-Improvement via Socratic-Guided Sampling
von: Ding, Yiwen, et al.
Veröffentlicht: (2024)
von: Ding, Yiwen, et al.
Veröffentlicht: (2024)
CodeChameleon: Personalized Encryption Framework for Jailbreaking Large Language Models
von: Lv, Huijie, et al.
Veröffentlicht: (2024)
von: Lv, Huijie, et al.
Veröffentlicht: (2024)
Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models
von: Wang, Shuting, et al.
Veröffentlicht: (2025)
von: Wang, Shuting, et al.
Veröffentlicht: (2025)
Outcome Accuracy is Not Enough: Aligning the Reasoning Process of Reward Models
von: Wang, Binghai, et al.
Veröffentlicht: (2026)
von: Wang, Binghai, et al.
Veröffentlicht: (2026)
Multi-Programming Language Sandbox for LLMs
von: Dou, Shihan, et al.
Veröffentlicht: (2024)
von: Dou, Shihan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
MetaRM: Shifted Distributions Alignment via Meta-Learning
von: Dou, Shihan, et al.
Veröffentlicht: (2024) -
StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback
von: Dou, Shihan, et al.
Veröffentlicht: (2024) -
Improving RL Exploration for LLM Reasoning through Retrospective Replay
von: Dou, Shihan, et al.
Veröffentlicht: (2025) -
Steering LLMs via Scalable Interactive Oversight
von: Zhou, Enyu, et al.
Veröffentlicht: (2026) -
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment
von: Zhang, Jiazheng, et al.
Veröffentlicht: (2025)