Ensembling Large Language Models with Process Reward-Guided Tree Search for Better Complex Reasoning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Park, Sungjin, Liu, Xiao, Gong, Yeyun, Choi, Edward
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912163280453632
author Park, Sungjin
Liu, Xiao
Gong, Yeyun
Choi, Edward
author_facet Park, Sungjin
Liu, Xiao
Gong, Yeyun
Choi, Edward
contents Despite recent advances in large language models, open-source models often struggle to consistently perform well on complex reasoning tasks. Existing ensemble methods, whether applied at the token or output levels, fail to address these challenges. In response, we present Language model Ensemble with Monte Carlo Tree Search (LE-MCTS), a novel framework for process-level ensembling of language models. LE-MCTS formulates step-by-step reasoning with an ensemble of language models as a Markov decision process. In this framework, states represent intermediate reasoning paths, while actions consist of generating the next reasoning step using one of the language models selected from a predefined pool. Guided by a process-based reward model, LE-MCTS performs a tree search over the reasoning steps generated by different language models, identifying the most accurate reasoning chain. Experimental results on five mathematical reasoning benchmarks demonstrate that our approach outperforms both single language model decoding algorithms and language model ensemble methods. Notably, LE-MCTS improves performance by 3.6% and 4.3% on the MATH and MQA datasets, respectively, highlighting its effectiveness in solving complex reasoning problems.
format Preprint
id arxiv_https___arxiv_org_abs_2412_15797
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Ensembling Large Language Models with Process Reward-Guided Tree Search for Better Complex Reasoning
Park, Sungjin
Liu, Xiao
Gong, Yeyun
Choi, Edward
Computation and Language
Despite recent advances in large language models, open-source models often struggle to consistently perform well on complex reasoning tasks. Existing ensemble methods, whether applied at the token or output levels, fail to address these challenges. In response, we present Language model Ensemble with Monte Carlo Tree Search (LE-MCTS), a novel framework for process-level ensembling of language models. LE-MCTS formulates step-by-step reasoning with an ensemble of language models as a Markov decision process. In this framework, states represent intermediate reasoning paths, while actions consist of generating the next reasoning step using one of the language models selected from a predefined pool. Guided by a process-based reward model, LE-MCTS performs a tree search over the reasoning steps generated by different language models, identifying the most accurate reasoning chain. Experimental results on five mathematical reasoning benchmarks demonstrate that our approach outperforms both single language model decoding algorithms and language model ensemble methods. Notably, LE-MCTS improves performance by 3.6% and 4.3% on the MATH and MQA datasets, respectively, highlighting its effectiveness in solving complex reasoning problems.
title Ensembling Large Language Models with Process Reward-Guided Tree Search for Better Complex Reasoning
topic Computation and Language
url https://arxiv.org/abs/2412.15797