Scaling over Scaling: Exploring Test-Time Scaling Plateau in Large Reasoning Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Jian, Zhu, Boyan, Leong, Chak Tou, Li, Yongqi, Li, Wenjie
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908398539243520
author Wang, Jian
Zhu, Boyan
Leong, Chak Tou
Li, Yongqi
Li, Wenjie
author_facet Wang, Jian
Zhu, Boyan
Leong, Chak Tou
Li, Yongqi
Li, Wenjie
contents Large reasoning models (LRMs) have exhibited the capacity of enhancing reasoning performance via internal test-time scaling. Building upon this, a promising direction is to further scale test-time compute to unlock even greater reasoning capabilities. However, as we push these scaling boundaries, systematically understanding the practical limits and achieving optimal resource allocation becomes a critical challenge. In this paper, we investigate the scaling plateau of test-time scaling and introduce the Test-Time Scaling Performance Model (TTSPM). We theoretically analyze two fundamental paradigms for such extended scaling, parallel scaling and sequential scaling, from a probabilistic modeling perspective. Our primary contribution is the derivation of the saturation point on the scaling budget for both strategies, identifying thresholds beyond which additional computation yields diminishing returns. Remarkably, despite their distinct mechanisms, both paradigms converge to a unified mathematical structure in their upper bounds. We empirically validate our theoretical findings on challenging reasoning benchmarks, including AIME, MATH-500, and GPQA, demonstrating the practical utility of these bounds for test-time resource allocation. We hope that this work provides insights into the cost-benefit trade-offs of test-time scaling, guiding the development of more resource-efficient inference strategies for large reasoning models.
format Preprint
id arxiv_https___arxiv_org_abs_2505_20522
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Scaling over Scaling: Exploring Test-Time Scaling Plateau in Large Reasoning Models
Wang, Jian
Zhu, Boyan
Leong, Chak Tou
Li, Yongqi
Li, Wenjie
Artificial Intelligence
Computation and Language
Machine Learning
Large reasoning models (LRMs) have exhibited the capacity of enhancing reasoning performance via internal test-time scaling. Building upon this, a promising direction is to further scale test-time compute to unlock even greater reasoning capabilities. However, as we push these scaling boundaries, systematically understanding the practical limits and achieving optimal resource allocation becomes a critical challenge. In this paper, we investigate the scaling plateau of test-time scaling and introduce the Test-Time Scaling Performance Model (TTSPM). We theoretically analyze two fundamental paradigms for such extended scaling, parallel scaling and sequential scaling, from a probabilistic modeling perspective. Our primary contribution is the derivation of the saturation point on the scaling budget for both strategies, identifying thresholds beyond which additional computation yields diminishing returns. Remarkably, despite their distinct mechanisms, both paradigms converge to a unified mathematical structure in their upper bounds. We empirically validate our theoretical findings on challenging reasoning benchmarks, including AIME, MATH-500, and GPQA, demonstrating the practical utility of these bounds for test-time resource allocation. We hope that this work provides insights into the cost-benefit trade-offs of test-time scaling, guiding the development of more resource-efficient inference strategies for large reasoning models.
title Scaling over Scaling: Exploring Test-Time Scaling Plateau in Large Reasoning Models
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2505.20522