Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Jiazheng, Jing, Wenqing, Zhang, Zizhuo, Xi, Zhiheng, Dou, Shihan, Weng, Rongxiang, Li, Jiahuan, Wang, Jingang, Chai, Mingxu, Hong, Shibo, Gui, Tao, Zhang, Qi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RMB: Comprehensively Benchmarking Reward Models in LLM Alignment
by: Zhou, Enyu, et al.
Published: (2024)
by: Zhou, Enyu, et al.
Published: (2024)
AgentV-RL: Scaling Reward Modeling with Agentic Verifier
by: Zhang, Jiazheng, et al.
Published: (2026)
by: Zhang, Jiazheng, et al.
Published: (2026)
Two Heads Are Better Than One: Dual-Model Verbal Reflection at Inference-Time
by: Li, Jiazheng, et al.
Published: (2025)
by: Li, Jiazheng, et al.
Published: (2025)
VRPO: Rethinking Value Modeling for Robust RL Training under Noisy Supervision
by: Zhu, Dingwei, et al.
Published: (2025)
by: Zhu, Dingwei, et al.
Published: (2025)
DocFusion: A Unified Framework for Document Parsing Tasks
by: Chai, Mingxu, et al.
Published: (2024)
by: Chai, Mingxu, et al.
Published: (2024)
Better Process Supervision with Bi-directional Rewarding Signals
by: Chen, Wenxiang, et al.
Published: (2025)
by: Chen, Wenxiang, et al.
Published: (2025)
Enhancing LLM-based Search Agents via Contribution Weighted Group Relative Policy Optimization
by: Wang, Junzhe, et al.
Published: (2026)
by: Wang, Junzhe, et al.
Published: (2026)
Two Heads Are Better Than One: Collaborative LLM Embodied Agents for Human-Robot Interaction
by: Rosser, Mitchell, et al.
Published: (2024)
by: Rosser, Mitchell, et al.
Published: (2024)
Can RL Improve Generalization of LLM Agents? An Empirical Study
by: Xi, Zhiheng, et al.
Published: (2026)
by: Xi, Zhiheng, et al.
Published: (2026)
DFPO: Scaling Value Modeling via Distributional Flow towards Robust and Generalizable LLM Post-Training
by: Zhu, Dingwei, et al.
Published: (2026)
by: Zhu, Dingwei, et al.
Published: (2026)
Two Heads are Better Than One: Test-time Scaling of Multi-agent Collaborative Reasoning
by: Jin, Can, et al.
Published: (2025)
by: Jin, Can, et al.
Published: (2025)
Prefix-Adaptive Block Diffusion for Efficient Document Recognition
by: Chai, Mingxu, et al.
Published: (2026)
by: Chai, Mingxu, et al.
Published: (2026)
Unlocking Implicit Experience: Synthesizing Tool-Use Trajectories from Text
by: Xu, Zhihao, et al.
Published: (2026)
by: Xu, Zhihao, et al.
Published: (2026)
Multi-Objective and Mixed-Reward Reinforcement Learning via Reward-Decorrelated Policy Optimization
by: Bai, Yang, et al.
Published: (2026)
by: Bai, Yang, et al.
Published: (2026)
Two Is Better Than One: Rotations Scale LoRAs
by: Guo, Hongcan, et al.
Published: (2025)
by: Guo, Hongcan, et al.
Published: (2025)
From Scores to Preferences: Redefining MOS Benchmarking for Speech Quality Reward Modeling
by: Cao, Yifei, et al.
Published: (2025)
by: Cao, Yifei, et al.
Published: (2025)
A Survey on LLM Mid-Training
by: Tu, Chengying, et al.
Published: (2025)
by: Tu, Chengying, et al.
Published: (2025)
Toward Optimal LLM Alignments Using Two-Player Games
by: Zheng, Rui, et al.
Published: (2024)
by: Zheng, Rui, et al.
Published: (2024)
Libra: Assessing and Improving Reward Model by Learning to Think
by: Zhou, Meng, et al.
Published: (2025)
by: Zhou, Meng, et al.
Published: (2025)
2Xplat: Two Experts Are Better Than One Generalist
by: Jeong, Hwasik, et al.
Published: (2026)
by: Jeong, Hwasik, et al.
Published: (2026)
Parental Presence at Induction—Are Two Parents Better Than One?
by: Nicole Almenrader, et al.
Published: (2025)
by: Nicole Almenrader, et al.
Published: (2025)
LLMEval-Med: A Real-world Clinical Benchmark for Medical LLMs with Physician Validation
by: Zhang, Ming, et al.
Published: (2025)
by: Zhang, Ming, et al.
Published: (2025)
DVPO: Distributional Value Modeling-based Policy Optimization for LLM Post-Training
by: Zhu, Dingwei, et al.
Published: (2025)
by: Zhu, Dingwei, et al.
Published: (2025)
Do Two Minds Work Better Than One? The Use of Cognitive Offloading to an External Agent in Working Memory Across Development
by: Chen Cheng
Published: (2026)
by: Chen Cheng
Published: (2026)
Phenomenology of Fractionally Charged Particles: Two Reps Are Better Than One
by: Koren, Seth, et al.
Published: (2025)
by: Koren, Seth, et al.
Published: (2025)
Two Is Better Than One: Aligned Representation Pairs for Anomaly Detection
by: Ryser, Alain, et al.
Published: (2024)
by: Ryser, Alain, et al.
Published: (2024)
Two Heads are Better Than One: Team Teaching in the Information Age.
by: Jurena, Donna Phin, et al.
Published: (1997)
by: Jurena, Donna Phin, et al.
Published: (1997)
JFTA-Bench: Evaluate LLM's Ability of Tracking and Analyzing Malfunctions Using Fault Trees
by: Wang, Yuhui, et al.
Published: (2026)
by: Wang, Yuhui, et al.
Published: (2026)
What's Wrong with Your Code Generated by Large Language Models? An Extensive Study
by: Dou, Shihan, et al.
Published: (2024)
by: Dou, Shihan, et al.
Published: (2024)
Two Intermediate Translations Are Better Than One: Fine-tuning LLMs for Document-level Translation Refinement
by: Dong, Yichen, et al.
Published: (2025)
by: Dong, Yichen, et al.
Published: (2025)
Two Heads Are Better Than One: Integrating Knowledge from Knowledge Graphs and Large Language Models for Entity Alignment
by: Yang, Linyao, et al.
Published: (2024)
by: Yang, Linyao, et al.
Published: (2024)
FIRE: Flexible Integration of Data Quality Ratings for Effective Pre-Training
by: Xu, Liangyu, et al.
Published: (2025)
by: Xu, Liangyu, et al.
Published: (2025)
LinkQA: Synthesizing Diverse QA from Multiple Seeds Strongly Linked by Knowledge Points
by: Zhang, Xuemiao, et al.
Published: (2025)
by: Zhang, Xuemiao, et al.
Published: (2025)
Large-Scale Diverse Synthesis for Mid-Training
by: Zhang, Xuemiao, et al.
Published: (2025)
by: Zhang, Xuemiao, et al.
Published: (2025)
Two Heads Are Better Than One: Boosting Graph Sparse Training via Semantic and Topological Awareness
by: Zhang, Guibin, et al.
Published: (2024)
by: Zhang, Guibin, et al.
Published: (2024)
Two is Better Than One: Digital Siblings to Improve Autonomous Driving Testing
by: Biagiola, Matteo, et al.
Published: (2023)
by: Biagiola, Matteo, et al.
Published: (2023)
When Are Two Scores Better Than One? Investigating Ensembles of Diffusion Models
by: Razafindralambo, Raphaël, et al.
Published: (2026)
by: Razafindralambo, Raphaël, et al.
Published: (2026)
Don't Squander Your Transition Year: Two Heads Are Better Than One
by: Trey Guinn, et al.
Published: (2025)
by: Trey Guinn, et al.
Published: (2025)
An Argument Against Committee Cochairs: Why One Leader Is Better Than Two
by: David C. Schwebel
Published: (2026)
by: David C. Schwebel
Published: (2026)
Improving RL Exploration for LLM Reasoning through Retrospective Replay
by: Dou, Shihan, et al.
Published: (2025)
by: Dou, Shihan, et al.
Published: (2025)
Similar Items
-
RMB: Comprehensively Benchmarking Reward Models in LLM Alignment
by: Zhou, Enyu, et al.
Published: (2024) -
AgentV-RL: Scaling Reward Modeling with Agentic Verifier
by: Zhang, Jiazheng, et al.
Published: (2026) -
Two Heads Are Better Than One: Dual-Model Verbal Reflection at Inference-Time
by: Li, Jiazheng, et al.
Published: (2025) -
VRPO: Rethinking Value Modeling for Robust RL Training under Noisy Supervision
by: Zhu, Dingwei, et al.
Published: (2025) -
DocFusion: A Unified Framework for Document Parsing Tasks
by: Chai, Mingxu, et al.
Published: (2024)