AceMath: Advancing Frontier Math Reasoning with Post-Training and Reward Modeling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Zihan, Chen, Yang, Shoeybi, Mohammad, Catanzaro, Bryan, Ping, Wei
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912191665405952
author Liu, Zihan
Chen, Yang
Shoeybi, Mohammad
Catanzaro, Bryan
Ping, Wei
author_facet Liu, Zihan
Chen, Yang
Shoeybi, Mohammad
Catanzaro, Bryan
Ping, Wei
contents In this paper, we introduce AceMath, a suite of frontier math models that excel in solving complex math problems, along with highly effective reward models capable of evaluating generated solutions and reliably identifying the correct ones. To develop the instruction-tuned math models, we propose a supervised fine-tuning (SFT) process that first achieves competitive performance across general domains, followed by targeted fine-tuning for the math domain using a carefully curated set of prompts and synthetically generated responses. The resulting model, AceMath-72B-Instruct greatly outperforms Qwen2.5-Math-72B-Instruct, GPT-4o and Claude-3.5 Sonnet. To develop math-specialized reward model, we first construct AceMath-RewardBench, a comprehensive and robust benchmark for evaluating math reward models across diverse problems and difficulty levels. After that, we present a systematic approach to build our math reward models. The resulting model, AceMath-72B-RM, consistently outperforms state-of-the-art reward models. Furthermore, when combining AceMath-72B-Instruct with AceMath-72B-RM, we achieve the highest average rm@8 score across the math reasoning benchmarks. We release model weights, training data, and evaluation benchmarks at: https://research.nvidia.com/labs/adlr/acemath
format Preprint
id arxiv_https___arxiv_org_abs_2412_15084
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle AceMath: Advancing Frontier Math Reasoning with Post-Training and Reward Modeling
Liu, Zihan
Chen, Yang
Shoeybi, Mohammad
Catanzaro, Bryan
Ping, Wei
Computation and Language
Artificial Intelligence
Machine Learning
In this paper, we introduce AceMath, a suite of frontier math models that excel in solving complex math problems, along with highly effective reward models capable of evaluating generated solutions and reliably identifying the correct ones. To develop the instruction-tuned math models, we propose a supervised fine-tuning (SFT) process that first achieves competitive performance across general domains, followed by targeted fine-tuning for the math domain using a carefully curated set of prompts and synthetically generated responses. The resulting model, AceMath-72B-Instruct greatly outperforms Qwen2.5-Math-72B-Instruct, GPT-4o and Claude-3.5 Sonnet. To develop math-specialized reward model, we first construct AceMath-RewardBench, a comprehensive and robust benchmark for evaluating math reward models across diverse problems and difficulty levels. After that, we present a systematic approach to build our math reward models. The resulting model, AceMath-72B-RM, consistently outperforms state-of-the-art reward models. Furthermore, when combining AceMath-72B-Instruct with AceMath-72B-RM, we achieve the highest average rm@8 score across the math reasoning benchmarks. We release model weights, training data, and evaluation benchmarks at: https://research.nvidia.com/labs/adlr/acemath
title AceMath: Advancing Frontier Math Reasoning with Post-Training and Reward Modeling
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2412.15084