Efficient LLM Collaboration via Planning

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Lee, Byeongchan, Lee, Jonghoon, Kim, Dongyoung, Kim, Jaehyung, Park, Kyungjoon, Lee, Dongjun, Shin, Jinwoo
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915999384600576
author Lee, Byeongchan
Lee, Jonghoon
Kim, Dongyoung
Kim, Jaehyung
Park, Kyungjoon
Lee, Dongjun
Shin, Jinwoo
author_facet Lee, Byeongchan
Lee, Jonghoon
Kim, Dongyoung
Kim, Jaehyung
Park, Kyungjoon
Lee, Dongjun
Shin, Jinwoo
contents Recently, large language models (LLMs) have demonstrated strong performance, ranging from simple to complex tasks. However, while large models achieve remarkable results across diverse tasks, they often incur substantial monetary inference cost, making frequent use impractical for many applications. In contrast, small models are often freely available and easy to deploy locally, but their performance on complex tasks remains limited. This trade-off raises a natural question: how can small and large models efficiently collaborate to combine their complementary strengths? To bridge this trade-off, we propose COPE, a test-time collaboration framework. A planner model first generates a plan that serves as a lightweight intermediate that guides a downstream executor model. Small and large models take turns acting as planner and executor, exchanging plans in a multi-stage cascade to collaboratively solve tasks. Through comprehensive experiments on benchmarks spanning mathematical reasoning, code generation, open-ended tasks, and agent tasks, we demonstrate that COPE achieves performance comparable to large proprietary models, while drastically reducing the inference API cost. These results highlight planning as an effective prior for cost-efficient inference.
format Preprint
id arxiv_https___arxiv_org_abs_2506_11578
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Efficient LLM Collaboration via Planning
Lee, Byeongchan
Lee, Jonghoon
Kim, Dongyoung
Kim, Jaehyung
Park, Kyungjoon
Lee, Dongjun
Shin, Jinwoo
Artificial Intelligence
Recently, large language models (LLMs) have demonstrated strong performance, ranging from simple to complex tasks. However, while large models achieve remarkable results across diverse tasks, they often incur substantial monetary inference cost, making frequent use impractical for many applications. In contrast, small models are often freely available and easy to deploy locally, but their performance on complex tasks remains limited. This trade-off raises a natural question: how can small and large models efficiently collaborate to combine their complementary strengths? To bridge this trade-off, we propose COPE, a test-time collaboration framework. A planner model first generates a plan that serves as a lightweight intermediate that guides a downstream executor model. Small and large models take turns acting as planner and executor, exchanging plans in a multi-stage cascade to collaboratively solve tasks. Through comprehensive experiments on benchmarks spanning mathematical reasoning, code generation, open-ended tasks, and agent tasks, we demonstrate that COPE achieves performance comparable to large proprietary models, while drastically reducing the inference API cost. These results highlight planning as an effective prior for cost-efficient inference.
title Efficient LLM Collaboration via Planning
topic Artificial Intelligence
url https://arxiv.org/abs/2506.11578