Parrot: A Training Pipeline Enhances Both Program CoT and Natural Language CoT for Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jin, Senjie, Chen, Lu, Xi, Zhiheng, Wang, Yuhui, Song, Sirui, Zhou, Yuhao, Zhang, Xinbo, Sun, Peng, Lu, Hong, Gui, Tao, Zhang, Qi, Huang, Xuanjing
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911239344488448
author Jin, Senjie
Chen, Lu
Xi, Zhiheng
Wang, Yuhui
Song, Sirui
Zhou, Yuhao
Zhang, Xinbo
Sun, Peng
Lu, Hong
Gui, Tao
Zhang, Qi
Huang, Xuanjing
author_facet Jin, Senjie
Chen, Lu
Xi, Zhiheng
Wang, Yuhui
Song, Sirui
Zhou, Yuhao
Zhang, Xinbo
Sun, Peng
Lu, Hong
Gui, Tao
Zhang, Qi
Huang, Xuanjing
contents Natural language chain-of-thought (N-CoT) and Program chain-of-thought (P-CoT) have emerged as two primary paradigms for large language models (LLMs) to solve mathematical reasoning problems. Current research typically endeavors to achieve unidirectional enhancement: P-CoT enhanced N-CoT or N-CoT enhanced P-CoT. In this paper, we seek to fully unleash the two paradigms' strengths for mutual enhancement and ultimately achieve simultaneous improvements. We conduct a detailed analysis of the error types across two paradigms, based on which we propose Parrot, a novel training pipeline for mathematical problems: 1) Three target-designed subtasks integrate sequential P-CoT and N-CoT generation. 2) A subtask hybrid training strategy to facilitate natural language semantic transferability. 3) The converted N-CoT auxiliary reward is designed to alleviate the sparse rewards in P-CoT optimization. Extensive experiments demonstrate that Parrot significantly enhances both the performance of N-CoT and P-CoT, especially on N-CoT. Using Parrot SFT, the N-CoT performance of LLaMA2 and CodeLLaMA achieve gains of +21.87 and +21.48 on MathQA over the RL baseline, which is resource-intensive.
format Preprint
id arxiv_https___arxiv_org_abs_2510_25310
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Parrot: A Training Pipeline Enhances Both Program CoT and Natural Language CoT for Reasoning
Jin, Senjie
Chen, Lu
Xi, Zhiheng
Wang, Yuhui
Song, Sirui
Zhou, Yuhao
Zhang, Xinbo
Sun, Peng
Lu, Hong
Gui, Tao
Zhang, Qi
Huang, Xuanjing
Computation and Language
Natural language chain-of-thought (N-CoT) and Program chain-of-thought (P-CoT) have emerged as two primary paradigms for large language models (LLMs) to solve mathematical reasoning problems. Current research typically endeavors to achieve unidirectional enhancement: P-CoT enhanced N-CoT or N-CoT enhanced P-CoT. In this paper, we seek to fully unleash the two paradigms' strengths for mutual enhancement and ultimately achieve simultaneous improvements. We conduct a detailed analysis of the error types across two paradigms, based on which we propose Parrot, a novel training pipeline for mathematical problems: 1) Three target-designed subtasks integrate sequential P-CoT and N-CoT generation. 2) A subtask hybrid training strategy to facilitate natural language semantic transferability. 3) The converted N-CoT auxiliary reward is designed to alleviate the sparse rewards in P-CoT optimization. Extensive experiments demonstrate that Parrot significantly enhances both the performance of N-CoT and P-CoT, especially on N-CoT. Using Parrot SFT, the N-CoT performance of LLaMA2 and CodeLLaMA achieve gains of +21.87 and +21.48 on MathQA over the RL baseline, which is resource-intensive.
title Parrot: A Training Pipeline Enhances Both Program CoT and Natural Language CoT for Reasoning
topic Computation and Language
url https://arxiv.org/abs/2510.25310