ReSo: A Reward-driven Self-organizing LLM-based Multi-Agent System for Reasoning Tasks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Heng, Geng, Hejia, Xue, Xiangyuan, Kang, Li, Qin, Yiran, Wang, Zhiyong, Yin, Zhenfei, Bai, Lei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915313439735808
author Zhou, Heng
Geng, Hejia
Xue, Xiangyuan
Kang, Li
Qin, Yiran
Wang, Zhiyong
Yin, Zhenfei
Bai, Lei
author_facet Zhou, Heng
Geng, Hejia
Xue, Xiangyuan
Kang, Li
Qin, Yiran
Wang, Zhiyong
Yin, Zhenfei
Bai, Lei
contents Multi-agent systems (MAS) have emerged as a promising approach for enhancing the reasoning capabilities of large language models in complex problem-solving; however, current MAS frameworks suffer from poor flexibility and scalability with underdeveloped optimization strategies. To address these challenges, we propose ReSo, which integrates task graph generation with a reward-driven two-stage agent selection process centered on our Collaborative Reward Model that provides fine-grained reward signals to optimize MAS cooperation. We also introduce an automated data synthesis framework for generating MAS benchmarks without any human annotations. Experimental results show that ReSo matches or outperforms existing methods, achieving 33.7 percent accuracy on Math-MAS and 32.3 percent accuracy on SciBench-MAS, where other approaches completely fail.
format Preprint
id arxiv_https___arxiv_org_abs_2503_02390
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ReSo: A Reward-driven Self-organizing LLM-based Multi-Agent System for Reasoning Tasks
Zhou, Heng
Geng, Hejia
Xue, Xiangyuan
Kang, Li
Qin, Yiran
Wang, Zhiyong
Yin, Zhenfei
Bai, Lei
Multiagent Systems
Multi-agent systems (MAS) have emerged as a promising approach for enhancing the reasoning capabilities of large language models in complex problem-solving; however, current MAS frameworks suffer from poor flexibility and scalability with underdeveloped optimization strategies. To address these challenges, we propose ReSo, which integrates task graph generation with a reward-driven two-stage agent selection process centered on our Collaborative Reward Model that provides fine-grained reward signals to optimize MAS cooperation. We also introduce an automated data synthesis framework for generating MAS benchmarks without any human annotations. Experimental results show that ReSo matches or outperforms existing methods, achieving 33.7 percent accuracy on Math-MAS and 32.3 percent accuracy on SciBench-MAS, where other approaches completely fail.
title ReSo: A Reward-driven Self-organizing LLM-based Multi-Agent System for Reasoning Tasks
topic Multiagent Systems
url https://arxiv.org/abs/2503.02390