MetaScale: Test-Time Scaling with Evolving Meta-Thoughts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Qin, Zhou, Wenxuan, Xu, Nan, Huang, James Y., Wang, Fei, Zhang, Sheng, Poon, Hoifung, Chen, Muhao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915202074673152
author Liu, Qin
Zhou, Wenxuan
Xu, Nan
Huang, James Y.
Wang, Fei
Zhang, Sheng
Poon, Hoifung
Chen, Muhao
author_facet Liu, Qin
Zhou, Wenxuan
Xu, Nan
Huang, James Y.
Wang, Fei
Zhang, Sheng
Poon, Hoifung
Chen, Muhao
contents One critical challenge for large language models (LLMs) for making complex reasoning is their reliance on matching reasoning patterns from training data, instead of proactively selecting the most appropriate cognitive strategy to solve a given task. Existing approaches impose fixed cognitive structures that enhance performance in specific tasks but lack adaptability across diverse scenarios. To address this limitation, we introduce METASCALE, a test-time scaling framework based on meta-thoughts -- adaptive thinking strategies tailored to each task. METASCALE initializes a pool of candidate meta-thoughts, then iteratively selects and evaluates them using a multi-armed bandit algorithm with upper confidence bound selection, guided by a reward model. To further enhance adaptability, a genetic algorithm evolves high-reward meta-thoughts, refining and extending the strategy pool over time. By dynamically proposing and optimizing meta-thoughts at inference time, METASCALE improves both accuracy and generalization across a wide range of tasks. Experimental results demonstrate that MetaScale consistently outperforms standard inference approaches, achieving an 11% performance gain in win rate on Arena-Hard for GPT-4o, surpassing o1-mini by 0.9% under style control. Notably, METASCALE scales more effectively with increasing sampling budgets and produces more structured, expert-level responses.
format Preprint
id arxiv_https___arxiv_org_abs_2503_13447
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MetaScale: Test-Time Scaling with Evolving Meta-Thoughts
Liu, Qin
Zhou, Wenxuan
Xu, Nan
Huang, James Y.
Wang, Fei
Zhang, Sheng
Poon, Hoifung
Chen, Muhao
Computation and Language
Artificial Intelligence
Machine Learning
One critical challenge for large language models (LLMs) for making complex reasoning is their reliance on matching reasoning patterns from training data, instead of proactively selecting the most appropriate cognitive strategy to solve a given task. Existing approaches impose fixed cognitive structures that enhance performance in specific tasks but lack adaptability across diverse scenarios. To address this limitation, we introduce METASCALE, a test-time scaling framework based on meta-thoughts -- adaptive thinking strategies tailored to each task. METASCALE initializes a pool of candidate meta-thoughts, then iteratively selects and evaluates them using a multi-armed bandit algorithm with upper confidence bound selection, guided by a reward model. To further enhance adaptability, a genetic algorithm evolves high-reward meta-thoughts, refining and extending the strategy pool over time. By dynamically proposing and optimizing meta-thoughts at inference time, METASCALE improves both accuracy and generalization across a wide range of tasks. Experimental results demonstrate that MetaScale consistently outperforms standard inference approaches, achieving an 11% performance gain in win rate on Arena-Hard for GPT-4o, surpassing o1-mini by 0.9% under style control. Notably, METASCALE scales more effectively with increasing sampling budgets and produces more structured, expert-level responses.
title MetaScale: Test-Time Scaling with Evolving Meta-Thoughts
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2503.13447