Thinking Long, but Short: Stable Sequential Test-Time Scaling for Large Reasoning Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Metel, Michael R., Cui, Yufei, Chen, Boxing, Parthasarathi, Prasanna
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908766841077760
author Metel, Michael R.
Cui, Yufei
Chen, Boxing
Parthasarathi, Prasanna
author_facet Metel, Michael R.
Cui, Yufei
Chen, Boxing
Parthasarathi, Prasanna
contents Sequential test-time scaling is a promising training-free method to improve large reasoning model accuracy, but as currently implemented, significant limitations have been observed. Inducing models to think for longer can increase their accuracy, but as the length of reasoning is further extended, it has also been shown to result in accuracy degradation and model instability. This work presents a novel sequential test-time scaling method, Min-Seek, which improves model accuracy significantly over a wide range of induced thoughts, stabilizing the accuracy of sequential scaling, and removing the need for reasoning length fine-tuning. Beyond improving model accuracy over a variety of reasoning tasks, our method is inherently efficient, as only the KV pairs of one additional induced thought are kept in the KV cache during reasoning. With a custom KV cache which stores keys without position embeddings, by dynamically encoding them contiguously before each new generated thought, our method can continue to reason well beyond a model's maximum context length, and under mild conditions has linear computational complexity.
format Preprint
id arxiv_https___arxiv_org_abs_2601_09855
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Thinking Long, but Short: Stable Sequential Test-Time Scaling for Large Reasoning Models
Metel, Michael R.
Cui, Yufei
Chen, Boxing
Parthasarathi, Prasanna
Artificial Intelligence
Computation and Language
Sequential test-time scaling is a promising training-free method to improve large reasoning model accuracy, but as currently implemented, significant limitations have been observed. Inducing models to think for longer can increase their accuracy, but as the length of reasoning is further extended, it has also been shown to result in accuracy degradation and model instability. This work presents a novel sequential test-time scaling method, Min-Seek, which improves model accuracy significantly over a wide range of induced thoughts, stabilizing the accuracy of sequential scaling, and removing the need for reasoning length fine-tuning. Beyond improving model accuracy over a variety of reasoning tasks, our method is inherently efficient, as only the KV pairs of one additional induced thought are kept in the KV cache during reasoning. With a custom KV cache which stores keys without position embeddings, by dynamically encoding them contiguously before each new generated thought, our method can continue to reason well beyond a model's maximum context length, and under mild conditions has linear computational complexity.
title Thinking Long, but Short: Stable Sequential Test-Time Scaling for Large Reasoning Models
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2601.09855