Step-level Value Preference Optimization for Mathematical Reasoning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Chen, Guoxin, Liao, Minpeng, Li, Chengxi, Fan, Kai
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916412242526208
author Chen, Guoxin
Liao, Minpeng
Li, Chengxi
Fan, Kai
author_facet Chen, Guoxin
Liao, Minpeng
Li, Chengxi
Fan, Kai
contents Direct Preference Optimization (DPO) using an implicit reward model has proven to be an effective alternative to reinforcement learning from human feedback (RLHF) for fine-tuning preference aligned large language models (LLMs). However, the overall preference annotations of responses do not fully capture the fine-grained quality of model outputs in complex multi-step reasoning tasks, such as mathematical reasoning. To address this limitation, we introduce a novel algorithm called Step-level Value Preference Optimization (SVPO). Our approach employs Monte Carlo Tree Search (MCTS) to automatically annotate step-level preferences for multi-step reasoning. Furthermore, from the perspective of learning-to-rank, we train an explicit value model to replicate the behavior of the implicit reward model, complementing standard preference optimization. This value model enables the LLM to generate higher reward responses with minimal cost during inference. Experimental results demonstrate that our method achieves state-of-the-art performance on both in-domain and out-of-domain mathematical reasoning benchmarks. Our code is available at \url{https://github.com/MARIO-Math-Reasoning/Super_MARIO}.
format Preprint
id arxiv_https___arxiv_org_abs_2406_10858
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Step-level Value Preference Optimization for Mathematical Reasoning
Chen, Guoxin
Liao, Minpeng
Li, Chengxi
Fan, Kai
Computation and Language
Artificial Intelligence
Direct Preference Optimization (DPO) using an implicit reward model has proven to be an effective alternative to reinforcement learning from human feedback (RLHF) for fine-tuning preference aligned large language models (LLMs). However, the overall preference annotations of responses do not fully capture the fine-grained quality of model outputs in complex multi-step reasoning tasks, such as mathematical reasoning. To address this limitation, we introduce a novel algorithm called Step-level Value Preference Optimization (SVPO). Our approach employs Monte Carlo Tree Search (MCTS) to automatically annotate step-level preferences for multi-step reasoning. Furthermore, from the perspective of learning-to-rank, we train an explicit value model to replicate the behavior of the implicit reward model, complementing standard preference optimization. This value model enables the LLM to generate higher reward responses with minimal cost during inference. Experimental results demonstrate that our method achieves state-of-the-art performance on both in-domain and out-of-domain mathematical reasoning benchmarks. Our code is available at \url{https://github.com/MARIO-Math-Reasoning/Super_MARIO}.
title Step-level Value Preference Optimization for Mathematical Reasoning
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2406.10858