A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Qiyuan, Lyu, Fuyuan, Sun, Zexu, Wang, Lei, Zhang, Weixu, Hua, Wenyue, Wu, Haolun, Guo, Zhihan, Wang, Yufei, Muennighoff, Niklas, King, Irwin, Liu, Xue, Ma, Chen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916719148138496
author Zhang, Qiyuan
Lyu, Fuyuan
Sun, Zexu
Wang, Lei
Zhang, Weixu
Hua, Wenyue
Wu, Haolun
Guo, Zhihan
Wang, Yufei
Muennighoff, Niklas
King, Irwin
Liu, Xue
Ma, Chen
author_facet Zhang, Qiyuan
Lyu, Fuyuan
Sun, Zexu
Wang, Lei
Zhang, Weixu
Hua, Wenyue
Wu, Haolun
Guo, Zhihan
Wang, Yufei
Muennighoff, Niklas
King, Irwin
Liu, Xue
Ma, Chen
contents As enthusiasm for scaling computation (data and parameters) in the pretraining era gradually diminished, test-time scaling (TTS), also referred to as ``test-time computing'' has emerged as a prominent research focus. Recent studies demonstrate that TTS can further elicit the problem-solving capabilities of large language models (LLMs), enabling significant breakthroughs not only in specialized reasoning tasks, such as mathematics and coding, but also in general tasks like open-ended Q&A. However, despite the explosion of recent efforts in this area, there remains an urgent need for a comprehensive survey offering a systemic understanding. To fill this gap, we propose a unified, multidimensional framework structured along four core dimensions of TTS research: what to scale, how to scale, where to scale, and how well to scale. Building upon this taxonomy, we conduct an extensive review of methods, application scenarios, and assessment aspects, and present an organized decomposition that highlights the unique functional roles of individual techniques within the broader TTS landscape. From this analysis, we distill the major developmental trajectories of TTS to date and offer hands-on guidelines for practical deployment. Furthermore, we identify several open challenges and offer insights into promising future directions, including further scaling, clarifying the functional essence of techniques, generalizing to more tasks, and more attributions. Our repository is available on https://github.com/testtimescaling/testtimescaling.github.io/
format Preprint
id arxiv_https___arxiv_org_abs_2503_24235
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?
Zhang, Qiyuan
Lyu, Fuyuan
Sun, Zexu
Wang, Lei
Zhang, Weixu
Hua, Wenyue
Wu, Haolun
Guo, Zhihan
Wang, Yufei
Muennighoff, Niklas
King, Irwin
Liu, Xue
Ma, Chen
Computation and Language
Artificial Intelligence
As enthusiasm for scaling computation (data and parameters) in the pretraining era gradually diminished, test-time scaling (TTS), also referred to as ``test-time computing'' has emerged as a prominent research focus. Recent studies demonstrate that TTS can further elicit the problem-solving capabilities of large language models (LLMs), enabling significant breakthroughs not only in specialized reasoning tasks, such as mathematics and coding, but also in general tasks like open-ended Q&A. However, despite the explosion of recent efforts in this area, there remains an urgent need for a comprehensive survey offering a systemic understanding. To fill this gap, we propose a unified, multidimensional framework structured along four core dimensions of TTS research: what to scale, how to scale, where to scale, and how well to scale. Building upon this taxonomy, we conduct an extensive review of methods, application scenarios, and assessment aspects, and present an organized decomposition that highlights the unique functional roles of individual techniques within the broader TTS landscape. From this analysis, we distill the major developmental trajectories of TTS to date and offer hands-on guidelines for practical deployment. Furthermore, we identify several open challenges and offer insights into promising future directions, including further scaling, clarifying the functional essence of techniques, generalizing to more tasks, and more attributions. Our repository is available on https://github.com/testtimescaling/testtimescaling.github.io/
title A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2503.24235