LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhang, Di, Wu, Jianbo, Lei, Jingdi, Che, Tong, Li, Jiatong, Xie, Tong, Huang, Xiaoshui, Zhang, Shufei, Pavone, Marco, Li, Yuqiang, Ouyang, Wanli, Zhou, Dongzhan
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917843170230272
author Zhang, Di
Wu, Jianbo
Lei, Jingdi
Che, Tong
Li, Jiatong
Xie, Tong
Huang, Xiaoshui
Zhang, Shufei
Pavone, Marco
Li, Yuqiang
Ouyang, Wanli
Zhou, Dongzhan
author_facet Zhang, Di
Wu, Jianbo
Lei, Jingdi
Che, Tong
Li, Jiatong
Xie, Tong
Huang, Xiaoshui
Zhang, Shufei
Pavone, Marco
Li, Yuqiang
Ouyang, Wanli
Zhou, Dongzhan
contents This paper presents an advanced mathematical problem-solving framework, LLaMA-Berry, for enhancing the mathematical reasoning ability of Large Language Models (LLMs). The framework combines Monte Carlo Tree Search (MCTS) with iterative Self-Refine to optimize the reasoning path and utilizes a pairwise reward model to evaluate different paths globally. By leveraging the self-critic and rewriting capabilities of LLMs, Self-Refine applied to MCTS (SR-MCTS) overcomes the inefficiencies and limitations of conventional step-wise and greedy search algorithms by fostering a more efficient exploration of solution spaces. Pairwise Preference Reward Model~(PPRM), inspired by Reinforcement Learning from Human Feedback (RLHF), is then used to model pairwise preferences between solutions, utilizing an Enhanced Borda Count (EBC) method to synthesize these preferences into a global ranking score to find better answers. This approach addresses the challenges of scoring variability and non-independent distributions in mathematical reasoning tasks. The framework has been tested on general and advanced benchmarks, showing superior performance in terms of search efficiency and problem-solving capability compared to existing methods like ToT and rStar, particularly in complex Olympiad-level benchmarks, including GPQA, AIME24 and AMC23.
format Preprint
id arxiv_https___arxiv_org_abs_2410_02884
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning
Zhang, Di
Wu, Jianbo
Lei, Jingdi
Che, Tong
Li, Jiatong
Xie, Tong
Huang, Xiaoshui
Zhang, Shufei
Pavone, Marco
Li, Yuqiang
Ouyang, Wanli
Zhou, Dongzhan
Artificial Intelligence
Computation and Language
This paper presents an advanced mathematical problem-solving framework, LLaMA-Berry, for enhancing the mathematical reasoning ability of Large Language Models (LLMs). The framework combines Monte Carlo Tree Search (MCTS) with iterative Self-Refine to optimize the reasoning path and utilizes a pairwise reward model to evaluate different paths globally. By leveraging the self-critic and rewriting capabilities of LLMs, Self-Refine applied to MCTS (SR-MCTS) overcomes the inefficiencies and limitations of conventional step-wise and greedy search algorithms by fostering a more efficient exploration of solution spaces. Pairwise Preference Reward Model~(PPRM), inspired by Reinforcement Learning from Human Feedback (RLHF), is then used to model pairwise preferences between solutions, utilizing an Enhanced Borda Count (EBC) method to synthesize these preferences into a global ranking score to find better answers. This approach addresses the challenges of scoring variability and non-independent distributions in mathematical reasoning tasks. The framework has been tested on general and advanced benchmarks, showing superior performance in terms of search efficiency and problem-solving capability compared to existing methods like ToT and rStar, particularly in complex Olympiad-level benchmarks, including GPQA, AIME24 and AMC23.
title LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2410.02884