Group Think: Multiple Concurrent Reasoning Agents Collaborating at Token Level Granularity

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hsu, Chan-Jan, Buffelli, Davide, McGowan, Jamie, Liao, Feng-Ting, Chen, Yi-Chang, Vakili, Sattar, Shiu, Da-shan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909612599410688
author Hsu, Chan-Jan
Buffelli, Davide
McGowan, Jamie
Liao, Feng-Ting
Chen, Yi-Chang
Vakili, Sattar
Shiu, Da-shan
author_facet Hsu, Chan-Jan
Buffelli, Davide
McGowan, Jamie
Liao, Feng-Ting
Chen, Yi-Chang
Vakili, Sattar
Shiu, Da-shan
contents Recent advances in large language models (LLMs) have demonstrated the power of reasoning through self-generated chains of thought. Multiple reasoning agents can collaborate to raise joint reasoning quality above individual outcomes. However, such agents typically interact in a turn-based manner, trading increased latency for improved quality. In this paper, we propose Group Think--a single LLM that acts as multiple concurrent reasoning agents, or thinkers. With shared visibility into each other's partial generation progress, Group Think introduces a new concurrent-reasoning paradigm in which multiple reasoning trajectories adapt dynamically to one another at the token level. For example, a reasoning thread may shift its generation mid-sentence upon detecting that another thread is better positioned to continue. This fine-grained, token-level collaboration enables Group Think to reduce redundant reasoning and improve quality while achieving significantly lower latency. Moreover, its concurrent nature allows for efficient utilization of idle computational resources, making it especially suitable for edge inference, where very small batch size often underutilizes local~GPUs. We give a simple and generalizable modification that enables any existing LLM to perform Group Think on a local GPU. We also present an evaluation strategy to benchmark reasoning latency and empirically demonstrate latency improvements using open-source LLMs that were not explicitly trained for Group Think. We hope this work paves the way for future LLMs to exhibit more sophisticated and more efficient collaborative behavior for higher quality generation.
format Preprint
id arxiv_https___arxiv_org_abs_2505_11107
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Group Think: Multiple Concurrent Reasoning Agents Collaborating at Token Level Granularity
Hsu, Chan-Jan
Buffelli, Davide
McGowan, Jamie
Liao, Feng-Ting
Chen, Yi-Chang
Vakili, Sattar
Shiu, Da-shan
Artificial Intelligence
Recent advances in large language models (LLMs) have demonstrated the power of reasoning through self-generated chains of thought. Multiple reasoning agents can collaborate to raise joint reasoning quality above individual outcomes. However, such agents typically interact in a turn-based manner, trading increased latency for improved quality. In this paper, we propose Group Think--a single LLM that acts as multiple concurrent reasoning agents, or thinkers. With shared visibility into each other's partial generation progress, Group Think introduces a new concurrent-reasoning paradigm in which multiple reasoning trajectories adapt dynamically to one another at the token level. For example, a reasoning thread may shift its generation mid-sentence upon detecting that another thread is better positioned to continue. This fine-grained, token-level collaboration enables Group Think to reduce redundant reasoning and improve quality while achieving significantly lower latency. Moreover, its concurrent nature allows for efficient utilization of idle computational resources, making it especially suitable for edge inference, where very small batch size often underutilizes local~GPUs. We give a simple and generalizable modification that enables any existing LLM to perform Group Think on a local GPU. We also present an evaluation strategy to benchmark reasoning latency and empirically demonstrate latency improvements using open-source LLMs that were not explicitly trained for Group Think. We hope this work paves the way for future LLMs to exhibit more sophisticated and more efficient collaborative behavior for higher quality generation.
title Group Think: Multiple Concurrent Reasoning Agents Collaborating at Token Level Granularity
topic Artificial Intelligence
url https://arxiv.org/abs/2505.11107