Enregistré dans:
Détails bibliographiques
Auteurs principaux: Dokmeci, Berkan, Wu, Qingyang, Athiwaratkun, Ben, Zhang, Ce, Song, Shuaiwen Leon, Zou, James
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:https://arxiv.org/abs/2507.02173
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909674391994368
author Dokmeci, Berkan
Wu, Qingyang
Athiwaratkun, Ben
Zhang, Ce
Song, Shuaiwen Leon
Zou, James
author_facet Dokmeci, Berkan
Wu, Qingyang
Athiwaratkun, Ben
Zhang, Ce
Song, Shuaiwen Leon
Zou, James
contents While recent advances in preference learning have enhanced alignment in human feedback, mathematical reasoning remains a persistent challenge. We investigate how data diversification strategies in preference optimization can improve the mathematical reasoning abilities of large language models (LLMs). We evaluate three common data generation methods: temperature sampling, Chain-of-Thought prompting, and Monte Carlo Tree Search (MCTS), and introduce Diversified-ThinkSolve (DTS), a novel structured approach that systematically decomposes problems into diverse reasoning paths. Our results show that with strategically diversified preference data, models can substantially improve mathematical reasoning performance, with the best approach yielding gains of 7.1% on GSM8K and 4.2% on MATH over the base model. Despite its strong performance, DTS incurs only a marginal computational overhead (1.03x) compared to the baseline, while MCTS is nearly five times more costly with lower returns. These findings demonstrate that structured exploration of diverse problem-solving methods creates more effective preference data for mathematical alignment than traditional approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2507_02173
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Data Diversification Methods In Alignment Enhance Math Performance In LLMs
Dokmeci, Berkan
Wu, Qingyang
Athiwaratkun, Ben
Zhang, Ce
Song, Shuaiwen Leon
Zou, James
Artificial Intelligence
While recent advances in preference learning have enhanced alignment in human feedback, mathematical reasoning remains a persistent challenge. We investigate how data diversification strategies in preference optimization can improve the mathematical reasoning abilities of large language models (LLMs). We evaluate three common data generation methods: temperature sampling, Chain-of-Thought prompting, and Monte Carlo Tree Search (MCTS), and introduce Diversified-ThinkSolve (DTS), a novel structured approach that systematically decomposes problems into diverse reasoning paths. Our results show that with strategically diversified preference data, models can substantially improve mathematical reasoning performance, with the best approach yielding gains of 7.1% on GSM8K and 4.2% on MATH over the base model. Despite its strong performance, DTS incurs only a marginal computational overhead (1.03x) compared to the baseline, while MCTS is nearly five times more costly with lower returns. These findings demonstrate that structured exploration of diverse problem-solving methods creates more effective preference data for mathematical alignment than traditional approaches.
title Data Diversification Methods In Alignment Enhance Math Performance In LLMs
topic Artificial Intelligence
url https://arxiv.org/abs/2507.02173