Every Rollout Counts: Optimal Resource Allocation for Efficient Test-Time Scaling

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Xinglin, Li, Yiwei, Feng, Shaoxiong, Yuan, Peiwen, Zhang, Yueqi, Shi, Jiayi, Tan, Chuyi, Pan, Boyuan, Hu, Yao, Li, Kan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911218963316736
author Wang, Xinglin
Li, Yiwei
Feng, Shaoxiong
Yuan, Peiwen
Zhang, Yueqi
Shi, Jiayi
Tan, Chuyi
Pan, Boyuan
Hu, Yao
Li, Kan
author_facet Wang, Xinglin
Li, Yiwei
Feng, Shaoxiong
Yuan, Peiwen
Zhang, Yueqi
Shi, Jiayi
Tan, Chuyi
Pan, Boyuan
Hu, Yao
Li, Kan
contents Test-Time Scaling (TTS) improves the performance of Large Language Models (LLMs) by using additional inference-time computation to explore multiple reasoning paths through search. Yet how to allocate a fixed rollout budget most effectively during search remains underexplored, often resulting in inefficient use of compute at test time. To bridge this gap, we formulate test-time search as a resource allocation problem and derive the optimal allocation strategy that maximizes the probability of obtaining a correct solution under a fixed rollout budget. Within this formulation, we reveal a core limitation of existing search methods: solution-level allocation tends to favor reasoning directions with more candidates, leading to theoretically suboptimal and inefficient use of compute. To address this, we propose Direction-Oriented Resource Allocation (DORA), a provably optimal method that mitigates this bias by decoupling direction quality from candidate count and allocating resources at the direction level. To demonstrate DORA's effectiveness, we conduct extensive experiments on challenging mathematical reasoning benchmarks including MATH500, AIME2024, and AIME2025. The empirical results show that DORA consistently outperforms strong baselines with comparable computational cost, achieving state-of-the-art accuracy. We hope our findings contribute to a broader understanding of optimal TTS for LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2506_15707
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Every Rollout Counts: Optimal Resource Allocation for Efficient Test-Time Scaling
Wang, Xinglin
Li, Yiwei
Feng, Shaoxiong
Yuan, Peiwen
Zhang, Yueqi
Shi, Jiayi
Tan, Chuyi
Pan, Boyuan
Hu, Yao
Li, Kan
Machine Learning
Artificial Intelligence
Test-Time Scaling (TTS) improves the performance of Large Language Models (LLMs) by using additional inference-time computation to explore multiple reasoning paths through search. Yet how to allocate a fixed rollout budget most effectively during search remains underexplored, often resulting in inefficient use of compute at test time. To bridge this gap, we formulate test-time search as a resource allocation problem and derive the optimal allocation strategy that maximizes the probability of obtaining a correct solution under a fixed rollout budget. Within this formulation, we reveal a core limitation of existing search methods: solution-level allocation tends to favor reasoning directions with more candidates, leading to theoretically suboptimal and inefficient use of compute. To address this, we propose Direction-Oriented Resource Allocation (DORA), a provably optimal method that mitigates this bias by decoupling direction quality from candidate count and allocating resources at the direction level. To demonstrate DORA's effectiveness, we conduct extensive experiments on challenging mathematical reasoning benchmarks including MATH500, AIME2024, and AIME2025. The empirical results show that DORA consistently outperforms strong baselines with comparable computational cost, achieving state-of-the-art accuracy. We hope our findings contribute to a broader understanding of optimal TTS for LLMs.
title Every Rollout Counts: Optimal Resource Allocation for Efficient Test-Time Scaling
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2506.15707