Understanding Performance Gap Between Parallel and Sequential Sampling in Large Reasoning Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Gu, Xiangming, De, Soham, Markeeva, Larisa, Veličković, Petar, Pascanu, Razvan
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917388354584576
author Gu, Xiangming
De, Soham
Markeeva, Larisa
Veličković, Petar
Pascanu, Razvan
author_facet Gu, Xiangming
De, Soham
Markeeva, Larisa
Veličković, Petar
Pascanu, Razvan
contents Large Reasoning Models (LRMs) have shown remarkable performance on challenging questions, such as math and coding. However, to obtain a high quality solution, one may need to sample more than once. In principal, there are two sampling strategies that can be composed to form more complex processes: sequential sampling and parallel sampling. In this paper, we first compare these two approaches with rigor, and observe, aligned with previous works, that parallel sampling seems to outperform sequential sampling even though the latter should have more representation power. To understand the underline reasons, we make three hypothesis on the reason behind this behavior: (i) parallel sampling outperforms due to the aggregator operator; (ii) sequential sampling is harmed by needing to use longer contexts; (iii) sequential sampling leads to less exploration due to conditioning on previous answers. The empirical evidence on various model families and sizes (Qwen3, DeepSeek-R1 distilled models, Gemini 2.5) and question domains (math and coding) suggests that the aggregation and context length do not seem to be the main culprit behind the performance gap. In contrast, the lack of exploration seems to play a considerably larger role, and we argue that this is one main cause for the performance gap.
format Preprint
id arxiv_https___arxiv_org_abs_2604_05868
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Understanding Performance Gap Between Parallel and Sequential Sampling in Large Reasoning Models
Gu, Xiangming
De, Soham
Markeeva, Larisa
Veličković, Petar
Pascanu, Razvan
Computation and Language
Large Reasoning Models (LRMs) have shown remarkable performance on challenging questions, such as math and coding. However, to obtain a high quality solution, one may need to sample more than once. In principal, there are two sampling strategies that can be composed to form more complex processes: sequential sampling and parallel sampling. In this paper, we first compare these two approaches with rigor, and observe, aligned with previous works, that parallel sampling seems to outperform sequential sampling even though the latter should have more representation power. To understand the underline reasons, we make three hypothesis on the reason behind this behavior: (i) parallel sampling outperforms due to the aggregator operator; (ii) sequential sampling is harmed by needing to use longer contexts; (iii) sequential sampling leads to less exploration due to conditioning on previous answers. The empirical evidence on various model families and sizes (Qwen3, DeepSeek-R1 distilled models, Gemini 2.5) and question domains (math and coding) suggests that the aggregation and context length do not seem to be the main culprit behind the performance gap. In contrast, the lack of exploration seems to play a considerably larger role, and we argue that this is one main cause for the performance gap.
title Understanding Performance Gap Between Parallel and Sequential Sampling in Large Reasoning Models
topic Computation and Language
url https://arxiv.org/abs/2604.05868