Test-Time Scaling in Reasoning Models Is Not Effective for Knowledge-Intensive Tasks Yet

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhao, James Xu, Hooi, Bryan, Ng, See-Kiong
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912865244413952
author Zhao, James Xu
Hooi, Bryan
Ng, See-Kiong
author_facet Zhao, James Xu
Hooi, Bryan
Ng, See-Kiong
contents Test-time scaling increases inference-time computation by allowing models to generate long reasoning chains, and has improved performance across many domains. However, in this work, we show that this approach is not yet effective for knowledge-intensive tasks. We evaluate 14 reasoning models on two knowledge-intensive benchmarks and find that increasing test-time computation does not consistently improve accuracy and often increases hallucinations. Further analysis shows that changes in hallucination rates under increased test-time computation are largely driven by models' willingness to answer. We also observe that extended reasoning can induce confirmation bias, leading to overconfident hallucinations. Finally, we provide an information-theoretic account: compute-only test-time scaling is a post-processing of a fixed trained model and therefore cannot increase information about the ground-truth answer beyond what is already encoded in the model, explaining its limited gains on knowledge-intensive tasks. Code and data are available at https://github.com/XuZhao0/tts-knowledge
format Preprint
id arxiv_https___arxiv_org_abs_2509_06861
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Test-Time Scaling in Reasoning Models Is Not Effective for Knowledge-Intensive Tasks Yet
Zhao, James Xu
Hooi, Bryan
Ng, See-Kiong
Artificial Intelligence
Computation and Language
Machine Learning
Test-time scaling increases inference-time computation by allowing models to generate long reasoning chains, and has improved performance across many domains. However, in this work, we show that this approach is not yet effective for knowledge-intensive tasks. We evaluate 14 reasoning models on two knowledge-intensive benchmarks and find that increasing test-time computation does not consistently improve accuracy and often increases hallucinations. Further analysis shows that changes in hallucination rates under increased test-time computation are largely driven by models' willingness to answer. We also observe that extended reasoning can induce confirmation bias, leading to overconfident hallucinations. Finally, we provide an information-theoretic account: compute-only test-time scaling is a post-processing of a fixed trained model and therefore cannot increase information about the ground-truth answer beyond what is already encoded in the model, explaining its limited gains on knowledge-intensive tasks. Code and data are available at https://github.com/XuZhao0/tts-knowledge
title Test-Time Scaling in Reasoning Models Is Not Effective for Knowledge-Intensive Tasks Yet
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2509.06861