IMAGINE: Integrating Multi-Agent System into One Model for Complex Reasoning and Planning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Xikai, Wang, Bo, Xiao, Likang, Li, Yongzhi, Chen, Quan, Wu, Wenjun, Liu, Liu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910017253277696
author Zhang, Xikai
Wang, Bo
Xiao, Likang
Li, Yongzhi
Chen, Quan
Wu, Wenjun
Liu, Liu
author_facet Zhang, Xikai
Wang, Bo
Xiao, Likang
Li, Yongzhi
Chen, Quan
Wu, Wenjun
Liu, Liu
contents Although large language models (LLMs) have made significant strides across various tasks, they still face significant challenges in complex reasoning and planning. For example, even with carefully designed prompts and prior information explicitly provided, GPT-4o achieves only a 7% Final Pass Rate on the TravelPlanner dataset in the sole-planning mode. Similarly, even in the thinking mode, Qwen3-8B-Instruct and DeepSeek-R1-671B, only achieve Final Pass Rates of 5.9% and 40%, respectively. Although well-organized Multi-Agent Systems (MAS) can offer improved collective reasoning, they often suffer from high reasoning costs due to multi-round internal interactions, long per-response latency, and difficulties in end-to-end training. To address these challenges, we propose a general and scalable framework called IMAGINE, short for Integrating Multi-Agent System into One Model. This framework not only integrates the reasoning and planning capabilities of MAS into a single, compact model, but also significantly surpass the capabilities of the MAS through a simple end-to-end training. Through this pipeline, a single small-scale model is not only able to acquire the structured reasoning and planning capabilities of a well-organized MAS but can also significantly outperform it. Experimental results demonstrate that, when using Qwen3-8B-Instruct as the base model and training it with our method, the model achieves an 82.7% Final Pass Rate on the TravelPlanner benchmark, far exceeding the 40% of DeepSeek-R1-671B, while maintaining a much smaller model size.
format Preprint
id arxiv_https___arxiv_org_abs_2510_14406
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle IMAGINE: Integrating Multi-Agent System into One Model for Complex Reasoning and Planning
Zhang, Xikai
Wang, Bo
Xiao, Likang
Li, Yongzhi
Chen, Quan
Wu, Wenjun
Liu, Liu
Artificial Intelligence
Computation and Language
Although large language models (LLMs) have made significant strides across various tasks, they still face significant challenges in complex reasoning and planning. For example, even with carefully designed prompts and prior information explicitly provided, GPT-4o achieves only a 7% Final Pass Rate on the TravelPlanner dataset in the sole-planning mode. Similarly, even in the thinking mode, Qwen3-8B-Instruct and DeepSeek-R1-671B, only achieve Final Pass Rates of 5.9% and 40%, respectively. Although well-organized Multi-Agent Systems (MAS) can offer improved collective reasoning, they often suffer from high reasoning costs due to multi-round internal interactions, long per-response latency, and difficulties in end-to-end training. To address these challenges, we propose a general and scalable framework called IMAGINE, short for Integrating Multi-Agent System into One Model. This framework not only integrates the reasoning and planning capabilities of MAS into a single, compact model, but also significantly surpass the capabilities of the MAS through a simple end-to-end training. Through this pipeline, a single small-scale model is not only able to acquire the structured reasoning and planning capabilities of a well-organized MAS but can also significantly outperform it. Experimental results demonstrate that, when using Qwen3-8B-Instruct as the base model and training it with our method, the model achieves an 82.7% Final Pass Rate on the TravelPlanner benchmark, far exceeding the 40% of DeepSeek-R1-671B, while maintaining a much smaller model size.
title IMAGINE: Integrating Multi-Agent System into One Model for Complex Reasoning and Planning
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2510.14406