PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909985636614144 |
|---|---|
| author | Hu, Jingcheng Zhang, Yinmin Shang, Shijie Yang, Xiaobo Peng, Yue Huang, Zhewei Zhou, Hebin Wu, Xin Cheng, Jie Wan, Fanqi Kong, Xiangwen Yao, Chengyuan Yan, Kaiwen Huang, Ailin Zhou, Hongyu Han, Qi Ge, Zheng Jiang, Daxin Zhang, Xiangyu Shum, Heung-Yeung |
| author_facet | Hu, Jingcheng Zhang, Yinmin Shang, Shijie Yang, Xiaobo Peng, Yue Huang, Zhewei Zhou, Hebin Wu, Xin Cheng, Jie Wan, Fanqi Kong, Xiangwen Yao, Chengyuan Yan, Kaiwen Huang, Ailin Zhou, Hongyu Han, Qi Ge, Zheng Jiang, Daxin Zhang, Xiangyu Shum, Heung-Yeung |
| contents | We introduce Parallel Coordinated Reasoning (PaCoRe), a training-and-inference framework designed to overcome a central limitation of contemporary language models: their inability to scale test-time compute (TTC) far beyond sequential reasoning under a fixed context window. PaCoRe departs from the traditional sequential paradigm by driving TTC through massive parallel exploration coordinated via a message-passing architecture in multiple rounds. Each round launches many parallel reasoning trajectories, compacts their findings into context-bounded messages, and synthesizes these messages to guide the next round and ultimately produce the final answer. Trained end-to-end with large-scale, outcome-based reinforcement learning, the model masters the synthesis abilities required by PaCoRe and scales to multi-million-token effective TTC without exceeding context limits. The approach yields strong improvements across diverse domains, and notably pushes reasoning beyond frontier systems in mathematics: an 8B model reaches 94.5% on HMMT 2025, surpassing GPT-5's 93.2% by scaling effective TTC to roughly two million tokens. We open-source model checkpoints, training data, and the full inference pipeline to accelerate follow-up work. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2601_05593 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning Hu, Jingcheng Zhang, Yinmin Shang, Shijie Yang, Xiaobo Peng, Yue Huang, Zhewei Zhou, Hebin Wu, Xin Cheng, Jie Wan, Fanqi Kong, Xiangwen Yao, Chengyuan Yan, Kaiwen Huang, Ailin Zhou, Hongyu Han, Qi Ge, Zheng Jiang, Daxin Zhang, Xiangyu Shum, Heung-Yeung Machine Learning We introduce Parallel Coordinated Reasoning (PaCoRe), a training-and-inference framework designed to overcome a central limitation of contemporary language models: their inability to scale test-time compute (TTC) far beyond sequential reasoning under a fixed context window. PaCoRe departs from the traditional sequential paradigm by driving TTC through massive parallel exploration coordinated via a message-passing architecture in multiple rounds. Each round launches many parallel reasoning trajectories, compacts their findings into context-bounded messages, and synthesizes these messages to guide the next round and ultimately produce the final answer. Trained end-to-end with large-scale, outcome-based reinforcement learning, the model masters the synthesis abilities required by PaCoRe and scales to multi-million-token effective TTC without exceeding context limits. The approach yields strong improvements across diverse domains, and notably pushes reasoning beyond frontier systems in mathematics: an 8B model reaches 94.5% on HMMT 2025, surpassing GPT-5's 93.2% by scaling effective TTC to roughly two million tokens. We open-source model checkpoints, training data, and the full inference pipeline to accelerate follow-up work. |
| title | PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2601.05593 |