Debate as Reward: A Multi-Agent Reward System for Scientific Ideation via RL Post-Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Salimi, Moein, Mohtasham, Babak Hosseini, Aghakasiri, Amin, Naieni, Mahdi, Qeysarbeigi, Amir Hossein, Nazer, Mohammad Masih Shalchian, Azar, Zahra, Siavoshani, Mahdi Jafari, Rohban, Mohammad Hossein
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913042038521856
author Salimi, Moein
Mohtasham, Babak Hosseini
Aghakasiri, Amin
Naieni, Mahdi
Qeysarbeigi, Amir Hossein
Nazer, Mohammad Masih Shalchian
Azar, Zahra
Siavoshani, Mahdi Jafari
Rohban, Mohammad Hossein
author_facet Salimi, Moein
Mohtasham, Babak Hosseini
Aghakasiri, Amin
Naieni, Mahdi
Qeysarbeigi, Amir Hossein
Nazer, Mohammad Masih Shalchian
Azar, Zahra
Siavoshani, Mahdi Jafari
Rohban, Mohammad Hossein
contents Large Language Models (LLMs) have demonstrated potential in automating scientific ideation, yet current approaches relying on iterative prompting or complex multi-agent architectures often suffer from hallucination or computational inefficiency. A critical bottleneck in applying Reinforcement Learning (RL) to this open-ended domain is reward hacking -- where models exploit imperfect evaluation proxies to maximize scores without producing genuine scientific innovation. To address these limitations, we propose an RL framework explicitly tailored for high-quality scientific idea generation. We propose the first multi-agent reward function designed to serve as a judge, decoupling methodological validation from implementation details while providing strict binary rewards that are robust to reward hacking. To effectively optimize against this sparse signal, we utilize an unbiased variant of Group Relative Policy Optimization to mitigate artificial length bias. We grounded our training in ICLR-320, a curated dataset of problem-solution pairs extracted from ICLR 2024 proceedings. Experiments demonstrate that our framework significantly outperforms state-of-the-art baselines across expert-evaluated metrics of novelty, feasibility, and effectiveness.
format Preprint
id arxiv_https___arxiv_org_abs_2604_16723
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Debate as Reward: A Multi-Agent Reward System for Scientific Ideation via RL Post-Training
Salimi, Moein
Mohtasham, Babak Hosseini
Aghakasiri, Amin
Naieni, Mahdi
Qeysarbeigi, Amir Hossein
Nazer, Mohammad Masih Shalchian
Azar, Zahra
Siavoshani, Mahdi Jafari
Rohban, Mohammad Hossein
Artificial Intelligence
Machine Learning
Large Language Models (LLMs) have demonstrated potential in automating scientific ideation, yet current approaches relying on iterative prompting or complex multi-agent architectures often suffer from hallucination or computational inefficiency. A critical bottleneck in applying Reinforcement Learning (RL) to this open-ended domain is reward hacking -- where models exploit imperfect evaluation proxies to maximize scores without producing genuine scientific innovation. To address these limitations, we propose an RL framework explicitly tailored for high-quality scientific idea generation. We propose the first multi-agent reward function designed to serve as a judge, decoupling methodological validation from implementation details while providing strict binary rewards that are robust to reward hacking. To effectively optimize against this sparse signal, we utilize an unbiased variant of Group Relative Policy Optimization to mitigate artificial length bias. We grounded our training in ICLR-320, a curated dataset of problem-solution pairs extracted from ICLR 2024 proceedings. Experiments demonstrate that our framework significantly outperforms state-of-the-art baselines across expert-evaluated metrics of novelty, feasibility, and effectiveness.
title Debate as Reward: A Multi-Agent Reward System for Scientific Ideation via RL Post-Training
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2604.16723