Agentic Aggregation for Parallel Scaling of Long-Horizon Agentic Tasks

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Lee, Yoonsang, Yen, Howard, Ye, Xi, Chen, Danqi
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908960312786944
author Lee, Yoonsang
Yen, Howard
Ye, Xi
Chen, Danqi
author_facet Lee, Yoonsang
Yen, Howard
Ye, Xi
Chen, Danqi
contents We study parallel test-time scaling for long-horizon agentic tasks such as agentic search and deep research, where multiple rollouts are generated in parallel and aggregated into a final response. While such scaling has proven effective for chain-of-thought reasoning, agentic tasks pose unique challenges: trajectories are long, multi-turn, and tool-augmented, and outputs are often open-ended. Aggregating only final answers discards rich information from trajectories, while concatenating all trajectories exceeds the model's context window. To address this, we propose AggAgent, an aggregation agent that treats parallel trajectories as an environment. We equip it with lightweight tools to inspect candidate solutions and search across trajectories, enabling it to navigate and synthesize information on demand. Across six benchmarks and three model families (GLM-4.7, Qwen3.5, MiniMax-M2.5), AggAgent outperforms all existing aggregation methods-by up to 5.3% absolute on average and 10.3% on two deep research tasks-while adding minimal overhead, as the aggregation cost remains bounded by a single agentic rollout. Our findings establish agentic aggregation as an effective and cost-efficient approach to parallel test-time scaling.
format Preprint
id arxiv_https___arxiv_org_abs_2604_11753
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Agentic Aggregation for Parallel Scaling of Long-Horizon Agentic Tasks
Lee, Yoonsang
Yen, Howard
Ye, Xi
Chen, Danqi
Computation and Language
We study parallel test-time scaling for long-horizon agentic tasks such as agentic search and deep research, where multiple rollouts are generated in parallel and aggregated into a final response. While such scaling has proven effective for chain-of-thought reasoning, agentic tasks pose unique challenges: trajectories are long, multi-turn, and tool-augmented, and outputs are often open-ended. Aggregating only final answers discards rich information from trajectories, while concatenating all trajectories exceeds the model's context window. To address this, we propose AggAgent, an aggregation agent that treats parallel trajectories as an environment. We equip it with lightweight tools to inspect candidate solutions and search across trajectories, enabling it to navigate and synthesize information on demand. Across six benchmarks and three model families (GLM-4.7, Qwen3.5, MiniMax-M2.5), AggAgent outperforms all existing aggregation methods-by up to 5.3% absolute on average and 10.3% on two deep research tasks-while adding minimal overhead, as the aggregation cost remains bounded by a single agentic rollout. Our findings establish agentic aggregation as an effective and cost-efficient approach to parallel test-time scaling.
title Agentic Aggregation for Parallel Scaling of Long-Horizon Agentic Tasks
topic Computation and Language
url https://arxiv.org/abs/2604.11753