Learning to Reason Across Parallel Samples for LLM Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qi, Jianing, Ye, Xi, Tang, Hao, Zhu, Zhigang, Choi, Eunsol
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912640056426496
author Qi, Jianing
Ye, Xi
Tang, Hao
Zhu, Zhigang
Choi, Eunsol
author_facet Qi, Jianing
Ye, Xi
Tang, Hao
Zhu, Zhigang
Choi, Eunsol
contents Scaling test-time compute brings substantial performance gains for large language models (LLMs). By sampling multiple answers and heuristically aggregate their answers (e.g., either through majority voting or using verifiers to rank the answers), one can achieve consistent performance gains in math domains. In this paper, we propose a new way to leverage such multiple sample set. We train a compact LLM, called Sample Set Aggregator (SSA), that takes a concatenated sequence of multiple samples and output the final answer, optimizing it for the answer accuracy with reinforcement learning. Experiments on five reasoning datasets demonstrate both the efficacy and efficiency of SSA. Notably, SSA improves over naive majority voting by 8% pass@5 on MATH. Furthermore, our 3B SSA surpasses model-based re-ranking with a much larger 72B process reward model. Our analysis also shows promising generalization ability of SSA, across sample set sizes, base model families and scales, and tasks. By separating LLMs to generate answers and LLMs to analyze and aggregate sampled answers, our approach can work with the outputs from premier black box models easily and efficiently.
format Preprint
id arxiv_https___arxiv_org_abs_2506_09014
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning to Reason Across Parallel Samples for LLM Reasoning
Qi, Jianing
Ye, Xi
Tang, Hao
Zhu, Zhigang
Choi, Eunsol
Computation and Language
Scaling test-time compute brings substantial performance gains for large language models (LLMs). By sampling multiple answers and heuristically aggregate their answers (e.g., either through majority voting or using verifiers to rank the answers), one can achieve consistent performance gains in math domains. In this paper, we propose a new way to leverage such multiple sample set. We train a compact LLM, called Sample Set Aggregator (SSA), that takes a concatenated sequence of multiple samples and output the final answer, optimizing it for the answer accuracy with reinforcement learning. Experiments on five reasoning datasets demonstrate both the efficacy and efficiency of SSA. Notably, SSA improves over naive majority voting by 8% pass@5 on MATH. Furthermore, our 3B SSA surpasses model-based re-ranking with a much larger 72B process reward model. Our analysis also shows promising generalization ability of SSA, across sample set sizes, base model families and scales, and tasks. By separating LLMs to generate answers and LLMs to analyze and aggregate sampled answers, our approach can work with the outputs from premier black box models easily and efficiently.
title Learning to Reason Across Parallel Samples for LLM Reasoning
topic Computation and Language
url https://arxiv.org/abs/2506.09014