TMAS: Scaling Test-Time Compute via Multi-Agent Synergy

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, George, Jing, Nan, Yi, Qing, Hao, Chuan, Yang, Ming, Chang, Feng, Wei, Yuan, Yang, Jian, Tao, Ran, Dai, Bryan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914578918539264
author Wu, George
Jing, Nan
Yi, Qing
Hao, Chuan
Yang, Ming
Chang, Feng
Wei, Yuan
Yang, Jian
Tao, Ran
Dai, Bryan
author_facet Wu, George
Jing, Nan
Yi, Qing
Hao, Chuan
Yang, Ming
Chang, Feng
Wei, Yuan
Yang, Jian
Tao, Ran
Dai, Bryan
contents Test-time scaling has become an effective paradigm for improving the reasoning ability of large language models by allocating additional computation during inference. Recent structured approaches have further advanced this paradigm by organizing inference across multiple trajectories, refinement rounds, and verification-based feedback. However, existing structured test-time scaling methods either weakly coordinate parallel reasoning trajectories or rely on noisy historical information without explicitly deciding what should be retained and reused, limiting their ability to balance exploration and exploitation. In this work, we propose TMAS, a framework for scaling test-time compute via multi-agent synergy. TMAS organizes inference as a collaborative process among specialized agents, enabling structured information flow across agents, trajectories, and refinement iterations. To support effective cross-trajectory collaboration, TMAS introduces hierarchical memories: the experience bank reuses low-level reliable intermediate conclusions and local feedback, while the guideline bank records previously explored high-level strategies to steer subsequent rollouts away from redundant reasoning patterns. Furthermore, we design a hybrid reward reinforcement learning scheme tailored to TMAS, which jointly preserves basic reasoning capability, enhances experience utilization, and encourages exploration beyond previously attempted solution strategies. Extensive experiments on challenging reasoning benchmarks show that TMAS achieves stronger iterative scaling than existing test-time scaling baselines, with hybrid reward training further improving scaling effectiveness and stability across iterations. Code and data are available at https://github.com/IQuestLab/tmas.
format Preprint
id arxiv_https___arxiv_org_abs_2605_10344
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle TMAS: Scaling Test-Time Compute via Multi-Agent Synergy
Wu, George
Jing, Nan
Yi, Qing
Hao, Chuan
Yang, Ming
Chang, Feng
Wei, Yuan
Yang, Jian
Tao, Ran
Dai, Bryan
Artificial Intelligence
Test-time scaling has become an effective paradigm for improving the reasoning ability of large language models by allocating additional computation during inference. Recent structured approaches have further advanced this paradigm by organizing inference across multiple trajectories, refinement rounds, and verification-based feedback. However, existing structured test-time scaling methods either weakly coordinate parallel reasoning trajectories or rely on noisy historical information without explicitly deciding what should be retained and reused, limiting their ability to balance exploration and exploitation. In this work, we propose TMAS, a framework for scaling test-time compute via multi-agent synergy. TMAS organizes inference as a collaborative process among specialized agents, enabling structured information flow across agents, trajectories, and refinement iterations. To support effective cross-trajectory collaboration, TMAS introduces hierarchical memories: the experience bank reuses low-level reliable intermediate conclusions and local feedback, while the guideline bank records previously explored high-level strategies to steer subsequent rollouts away from redundant reasoning patterns. Furthermore, we design a hybrid reward reinforcement learning scheme tailored to TMAS, which jointly preserves basic reasoning capability, enhances experience utilization, and encourages exploration beyond previously attempted solution strategies. Extensive experiments on challenging reasoning benchmarks show that TMAS achieves stronger iterative scaling than existing test-time scaling baselines, with hybrid reward training further improving scaling effectiveness and stability across iterations. Code and data are available at https://github.com/IQuestLab/tmas.
title TMAS: Scaling Test-Time Compute via Multi-Agent Synergy
topic Artificial Intelligence
url https://arxiv.org/abs/2605.10344