A2R: An Asymmetric Two-Stage Reasoning Framework for Parallel Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Ziqi, Niu, Boye, Li, Zhongli, Meng, Linghui, Liu, Jing, Zheng, Zhi, Xu, Tong, Wu, Hua, Wang, Haifeng, Chen, Enhong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916971652579328
author Wang, Ziqi
Niu, Boye
Li, Zhongli
Meng, Linghui
Liu, Jing
Zheng, Zhi
Xu, Tong
Wu, Hua
Wang, Haifeng
Chen, Enhong
author_facet Wang, Ziqi
Niu, Boye
Li, Zhongli
Meng, Linghui
Liu, Jing
Zheng, Zhi
Xu, Tong
Wu, Hua
Wang, Haifeng
Chen, Enhong
contents Recent Large Reasoning Models have achieved significant improvements in complex task-solving capabilities by allocating more computation at the inference stage with a "thinking longer" paradigm. Even as the foundational reasoning capabilities of models advance rapidly, the persistent gap between a model's performance in a single attempt and its latent potential, often revealed only across multiple solution paths, starkly highlights the disparity between its realized and inherent capabilities. To address this, we present A2R, an Asymmetric Two-Stage Reasoning framework designed to explicitly bridge the gap between a model's potential and its actual performance. In this framework, an "explorer" model first generates potential solutions in parallel through repeated sampling. Subsequently,a "synthesizer" model integrates these references for a more refined, second stage of reasoning. This two-stage process allows computation to be scaled orthogonally to existing sequential methods. Our work makes two key innovations: First, we present A2R as a plug-and-play parallel reasoning framework that explicitly enhances a model's capabilities on complex questions. For example, using our framework, the Qwen3-8B-distill model achieves a 75% performance improvement compared to its self-consistency baseline. Second, through a systematic analysis of the explorer and synthesizer roles, we identify an effective asymmetric scaling paradigm. This insight leads to A2R-Efficient, a "small-to-big" variant that combines a Qwen3-4B explorer with a Qwen3-8B synthesizer. This configuration surpasses the average performance of a monolithic Qwen3-32B model at a nearly 30% lower cost. Collectively, these results show that A2R is not only a performance-boosting framework but also an efficient and practical solution for real-world applications.
format Preprint
id arxiv_https___arxiv_org_abs_2509_22044
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A2R: An Asymmetric Two-Stage Reasoning Framework for Parallel Reasoning
Wang, Ziqi
Niu, Boye
Li, Zhongli
Meng, Linghui
Liu, Jing
Zheng, Zhi
Xu, Tong
Wu, Hua
Wang, Haifeng
Chen, Enhong
Artificial Intelligence
Computation and Language
Recent Large Reasoning Models have achieved significant improvements in complex task-solving capabilities by allocating more computation at the inference stage with a "thinking longer" paradigm. Even as the foundational reasoning capabilities of models advance rapidly, the persistent gap between a model's performance in a single attempt and its latent potential, often revealed only across multiple solution paths, starkly highlights the disparity between its realized and inherent capabilities. To address this, we present A2R, an Asymmetric Two-Stage Reasoning framework designed to explicitly bridge the gap between a model's potential and its actual performance. In this framework, an "explorer" model first generates potential solutions in parallel through repeated sampling. Subsequently,a "synthesizer" model integrates these references for a more refined, second stage of reasoning. This two-stage process allows computation to be scaled orthogonally to existing sequential methods. Our work makes two key innovations: First, we present A2R as a plug-and-play parallel reasoning framework that explicitly enhances a model's capabilities on complex questions. For example, using our framework, the Qwen3-8B-distill model achieves a 75% performance improvement compared to its self-consistency baseline. Second, through a systematic analysis of the explorer and synthesizer roles, we identify an effective asymmetric scaling paradigm. This insight leads to A2R-Efficient, a "small-to-big" variant that combines a Qwen3-4B explorer with a Qwen3-8B synthesizer. This configuration surpasses the average performance of a monolithic Qwen3-32B model at a nearly 30% lower cost. Collectively, these results show that A2R is not only a performance-boosting framework but also an efficient and practical solution for real-world applications.
title A2R: An Asymmetric Two-Stage Reasoning Framework for Parallel Reasoning
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2509.22044