ReSo: A Reward-driven Self-organizing LLM-based Multi-Agent System for Reasoning Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915313439735808 |
|---|---|
| author | Zhou, Heng Geng, Hejia Xue, Xiangyuan Kang, Li Qin, Yiran Wang, Zhiyong Yin, Zhenfei Bai, Lei |
| author_facet | Zhou, Heng Geng, Hejia Xue, Xiangyuan Kang, Li Qin, Yiran Wang, Zhiyong Yin, Zhenfei Bai, Lei |
| contents | Multi-agent systems (MAS) have emerged as a promising approach for enhancing the reasoning capabilities of large language models in complex problem-solving; however, current MAS frameworks suffer from poor flexibility and scalability with underdeveloped optimization strategies. To address these challenges, we propose ReSo, which integrates task graph generation with a reward-driven two-stage agent selection process centered on our Collaborative Reward Model that provides fine-grained reward signals to optimize MAS cooperation. We also introduce an automated data synthesis framework for generating MAS benchmarks without any human annotations. Experimental results show that ReSo matches or outperforms existing methods, achieving 33.7 percent accuracy on Math-MAS and 32.3 percent accuracy on SciBench-MAS, where other approaches completely fail. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_02390 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | ReSo: A Reward-driven Self-organizing LLM-based Multi-Agent System for Reasoning Tasks Zhou, Heng Geng, Hejia Xue, Xiangyuan Kang, Li Qin, Yiran Wang, Zhiyong Yin, Zhenfei Bai, Lei Multiagent Systems Multi-agent systems (MAS) have emerged as a promising approach for enhancing the reasoning capabilities of large language models in complex problem-solving; however, current MAS frameworks suffer from poor flexibility and scalability with underdeveloped optimization strategies. To address these challenges, we propose ReSo, which integrates task graph generation with a reward-driven two-stage agent selection process centered on our Collaborative Reward Model that provides fine-grained reward signals to optimize MAS cooperation. We also introduce an automated data synthesis framework for generating MAS benchmarks without any human annotations. Experimental results show that ReSo matches or outperforms existing methods, achieving 33.7 percent accuracy on Math-MAS and 32.3 percent accuracy on SciBench-MAS, where other approaches completely fail. |
| title | ReSo: A Reward-driven Self-organizing LLM-based Multi-Agent System for Reasoning Tasks |
| topic | Multiagent Systems |
| url | https://arxiv.org/abs/2503.02390 |