ThetaEvolve: Test-time Learning on Open Problems

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Yiping, Su, Shao-Rong, Zeng, Zhiyuan, Xu, Eva, Ren, Liliang, Yang, Xinyu, Huang, Zeyi, He, Xuehai, Ma, Luyao, Peng, Baolin, Cheng, Hao, He, Pengcheng, Chen, Weizhu, Wang, Shuohang, Du, Simon Shaolei, Shen, Yelong
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911292726444032
author Wang, Yiping
Su, Shao-Rong
Zeng, Zhiyuan
Xu, Eva
Ren, Liliang
Yang, Xinyu
Huang, Zeyi
He, Xuehai
Ma, Luyao
Peng, Baolin
Cheng, Hao
He, Pengcheng
Chen, Weizhu
Wang, Shuohang
Du, Simon Shaolei
Shen, Yelong
author_facet Wang, Yiping
Su, Shao-Rong
Zeng, Zhiyuan
Xu, Eva
Ren, Liliang
Yang, Xinyu
Huang, Zeyi
He, Xuehai
Ma, Luyao
Peng, Baolin
Cheng, Hao
He, Pengcheng
Chen, Weizhu
Wang, Shuohang
Du, Simon Shaolei
Shen, Yelong
contents Recent advances in large language models (LLMs) have enabled breakthroughs in mathematical discovery, exemplified by AlphaEvolve, a closed-source system that evolves programs to improve bounds on open problems. However, it relies on ensembles of frontier LLMs to achieve new bounds and is a pure inference system that models cannot internalize the evolving strategies. We introduce ThetaEvolve, an open-source framework that simplifies and extends AlphaEvolve to efficiently scale both in-context learning and Reinforcement Learning (RL) at test time, allowing models to continually learn from their experiences in improving open optimization problems. ThetaEvolve features a single LLM, a large program database for enhanced exploration, batch sampling for higher throughput, lazy penalties to discourage stagnant outputs, and optional reward shaping for stable training signals, etc. ThetaEvolve is the first evolving framework that enable a small open-source model, like DeepSeek-R1-0528-Qwen3-8B, to achieve new best-known bounds on open problems (circle packing and first auto-correlation inequality) mentioned in AlphaEvolve. Besides, across two models and four open tasks, we find that ThetaEvolve with RL at test-time consistently outperforms inference-only baselines, and the model indeed learns evolving capabilities, as the RL-trained checkpoints demonstrate faster progress and better final performance on both trained target task and other unseen tasks. We release our code publicly: https://github.com/ypwang61/ThetaEvolve
format Preprint
id arxiv_https___arxiv_org_abs_2511_23473
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ThetaEvolve: Test-time Learning on Open Problems
Wang, Yiping
Su, Shao-Rong
Zeng, Zhiyuan
Xu, Eva
Ren, Liliang
Yang, Xinyu
Huang, Zeyi
He, Xuehai
Ma, Luyao
Peng, Baolin
Cheng, Hao
He, Pengcheng
Chen, Weizhu
Wang, Shuohang
Du, Simon Shaolei
Shen, Yelong
Machine Learning
Computation and Language
Recent advances in large language models (LLMs) have enabled breakthroughs in mathematical discovery, exemplified by AlphaEvolve, a closed-source system that evolves programs to improve bounds on open problems. However, it relies on ensembles of frontier LLMs to achieve new bounds and is a pure inference system that models cannot internalize the evolving strategies. We introduce ThetaEvolve, an open-source framework that simplifies and extends AlphaEvolve to efficiently scale both in-context learning and Reinforcement Learning (RL) at test time, allowing models to continually learn from their experiences in improving open optimization problems. ThetaEvolve features a single LLM, a large program database for enhanced exploration, batch sampling for higher throughput, lazy penalties to discourage stagnant outputs, and optional reward shaping for stable training signals, etc. ThetaEvolve is the first evolving framework that enable a small open-source model, like DeepSeek-R1-0528-Qwen3-8B, to achieve new best-known bounds on open problems (circle packing and first auto-correlation inequality) mentioned in AlphaEvolve. Besides, across two models and four open tasks, we find that ThetaEvolve with RL at test-time consistently outperforms inference-only baselines, and the model indeed learns evolving capabilities, as the RL-trained checkpoints demonstrate faster progress and better final performance on both trained target task and other unseen tasks. We release our code publicly: https://github.com/ypwang61/ThetaEvolve
title ThetaEvolve: Test-time Learning on Open Problems
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2511.23473