Salvato in:
Dettagli Bibliografici
Autori principali: Chen, Hongwei, Lei, Yishu, Zhang, Dan, Ke, Bo, Zhu, Danxiang, Chen, Xuyi, Lu, Yuxiang, Huang, Zhengjie, Feng, Shikun, He, Jingzhou, Sun, Yu, Wu, Hua, Wang, Haifeng
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:https://arxiv.org/abs/2510.10293
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917005317111808
author Chen, Hongwei
Lei, Yishu
Zhang, Dan
Ke, Bo
Zhu, Danxiang
Chen, Xuyi
Lu, Yuxiang
Huang, Zhengjie
Feng, Shikun
He, Jingzhou
Sun, Yu
Wu, Hua
Wang, Haifeng
author_facet Chen, Hongwei
Lei, Yishu
Zhang, Dan
Ke, Bo
Zhu, Danxiang
Chen, Xuyi
Lu, Yuxiang
Huang, Zhengjie
Feng, Shikun
He, Jingzhou
Sun, Yu
Wu, Hua
Wang, Haifeng
contents Test-time scaling has emerged as a promising paradigm in language modeling, wherein additional computational resources are allocated during inference to enhance model performance. Recent approaches, such as DeepConf, have demonstrated the efficacy of this strategy, however, they often incur substantial computational overhead to achieve competitive results. In this work, we propose MatryoshkaThinking, a novel method that significantly reduces computational cost while maintaining state-of-the-art performance. Specifically, MatryoshkaThinking attains a score of 99.79 on AIME2025 using only 4% of the computation required by DeepConf. The core of our approach lies in the recursive exploitation of the model's intrinsic capabilities in reasoning, verification, and summarization, which collectively enhance the retention of correct solutions and reduce the disparity between Pass@k and Pass@1. Comprehensive evaluations across multiple open-source models and challenging multi-modal reasoning benchmarks validate the effectiveness and generality of our method. These findings offer new insights into the design of efficient and scalable test-time inference strategies for advanced language models.
format Preprint
id arxiv_https___arxiv_org_abs_2510_10293
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MatryoshkaThinking: Recursive Test-Time Scaling Enables Efficient Reasoning
Chen, Hongwei
Lei, Yishu
Zhang, Dan
Ke, Bo
Zhu, Danxiang
Chen, Xuyi
Lu, Yuxiang
Huang, Zhengjie
Feng, Shikun
He, Jingzhou
Sun, Yu
Wu, Hua
Wang, Haifeng
Computation and Language
Artificial Intelligence
Test-time scaling has emerged as a promising paradigm in language modeling, wherein additional computational resources are allocated during inference to enhance model performance. Recent approaches, such as DeepConf, have demonstrated the efficacy of this strategy, however, they often incur substantial computational overhead to achieve competitive results. In this work, we propose MatryoshkaThinking, a novel method that significantly reduces computational cost while maintaining state-of-the-art performance. Specifically, MatryoshkaThinking attains a score of 99.79 on AIME2025 using only 4% of the computation required by DeepConf. The core of our approach lies in the recursive exploitation of the model's intrinsic capabilities in reasoning, verification, and summarization, which collectively enhance the retention of correct solutions and reduce the disparity between Pass@k and Pass@1. Comprehensive evaluations across multiple open-source models and challenging multi-modal reasoning benchmarks validate the effectiveness and generality of our method. These findings offer new insights into the design of efficient and scalable test-time inference strategies for advanced language models.
title MatryoshkaThinking: Recursive Test-Time Scaling Enables Efficient Reasoning
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.10293