PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Jingcheng, Zhang, Yinmin, Shang, Shijie, Yang, Xiaobo, Peng, Yue, Huang, Zhewei, Zhou, Hebin, Wu, Xin, Cheng, Jie, Wan, Fanqi, Kong, Xiangwen, Yao, Chengyuan, Yan, Kaiwen, Huang, Ailin, Zhou, Hongyu, Han, Qi, Ge, Zheng, Jiang, Daxin, Zhang, Xiangyu, Shum, Heung-Yeung
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909985636614144
author Hu, Jingcheng
Zhang, Yinmin
Shang, Shijie
Yang, Xiaobo
Peng, Yue
Huang, Zhewei
Zhou, Hebin
Wu, Xin
Cheng, Jie
Wan, Fanqi
Kong, Xiangwen
Yao, Chengyuan
Yan, Kaiwen
Huang, Ailin
Zhou, Hongyu
Han, Qi
Ge, Zheng
Jiang, Daxin
Zhang, Xiangyu
Shum, Heung-Yeung
author_facet Hu, Jingcheng
Zhang, Yinmin
Shang, Shijie
Yang, Xiaobo
Peng, Yue
Huang, Zhewei
Zhou, Hebin
Wu, Xin
Cheng, Jie
Wan, Fanqi
Kong, Xiangwen
Yao, Chengyuan
Yan, Kaiwen
Huang, Ailin
Zhou, Hongyu
Han, Qi
Ge, Zheng
Jiang, Daxin
Zhang, Xiangyu
Shum, Heung-Yeung
contents We introduce Parallel Coordinated Reasoning (PaCoRe), a training-and-inference framework designed to overcome a central limitation of contemporary language models: their inability to scale test-time compute (TTC) far beyond sequential reasoning under a fixed context window. PaCoRe departs from the traditional sequential paradigm by driving TTC through massive parallel exploration coordinated via a message-passing architecture in multiple rounds. Each round launches many parallel reasoning trajectories, compacts their findings into context-bounded messages, and synthesizes these messages to guide the next round and ultimately produce the final answer. Trained end-to-end with large-scale, outcome-based reinforcement learning, the model masters the synthesis abilities required by PaCoRe and scales to multi-million-token effective TTC without exceeding context limits. The approach yields strong improvements across diverse domains, and notably pushes reasoning beyond frontier systems in mathematics: an 8B model reaches 94.5% on HMMT 2025, surpassing GPT-5's 93.2% by scaling effective TTC to roughly two million tokens. We open-source model checkpoints, training data, and the full inference pipeline to accelerate follow-up work.
format Preprint
id arxiv_https___arxiv_org_abs_2601_05593
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning
Hu, Jingcheng
Zhang, Yinmin
Shang, Shijie
Yang, Xiaobo
Peng, Yue
Huang, Zhewei
Zhou, Hebin
Wu, Xin
Cheng, Jie
Wan, Fanqi
Kong, Xiangwen
Yao, Chengyuan
Yan, Kaiwen
Huang, Ailin
Zhou, Hongyu
Han, Qi
Ge, Zheng
Jiang, Daxin
Zhang, Xiangyu
Shum, Heung-Yeung
Machine Learning
We introduce Parallel Coordinated Reasoning (PaCoRe), a training-and-inference framework designed to overcome a central limitation of contemporary language models: their inability to scale test-time compute (TTC) far beyond sequential reasoning under a fixed context window. PaCoRe departs from the traditional sequential paradigm by driving TTC through massive parallel exploration coordinated via a message-passing architecture in multiple rounds. Each round launches many parallel reasoning trajectories, compacts their findings into context-bounded messages, and synthesizes these messages to guide the next round and ultimately produce the final answer. Trained end-to-end with large-scale, outcome-based reinforcement learning, the model masters the synthesis abilities required by PaCoRe and scales to multi-million-token effective TTC without exceeding context limits. The approach yields strong improvements across diverse domains, and notably pushes reasoning beyond frontier systems in mathematics: an 8B model reaches 94.5% on HMMT 2025, surpassing GPT-5's 93.2% by scaling effective TTC to roughly two million tokens. We open-source model checkpoints, training data, and the full inference pipeline to accelerate follow-up work.
title PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning
topic Machine Learning
url https://arxiv.org/abs/2601.05593