Exploring Mathematical Extrapolation of Large Language Models with Synthetic Data

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Li, Haolong, Ma, Yu, Zhang, Yinqi, Ye, Chen, Chen, Jie
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911903258771456
author Li, Haolong
Ma, Yu
Zhang, Yinqi
Ye, Chen
Chen, Jie
author_facet Li, Haolong
Ma, Yu
Zhang, Yinqi
Ye, Chen
Chen, Jie
contents Large Language Models (LLMs) have shown excellent performance in language understanding, text generation, code synthesis, and many other tasks, while they still struggle in complex multi-step reasoning problems, such as mathematical reasoning. In this paper, through a newly proposed arithmetical puzzle problem, we show that the model can perform well on multi-step reasoning tasks via fine-tuning on high-quality synthetic data. Experimental results with the open-llama-3B model on three different test datasets show that not only the model can reach a zero-shot pass@1 at 0.44 on the in-domain dataset, it also demonstrates certain generalization capabilities on the out-of-domain datasets. Specifically, this paper has designed two out-of-domain datasets in the form of extending the numerical range and the composing components of the arithmetical puzzle problem separately. The fine-tuned models have shown encouraging performance on these two far more difficult tasks with the zero-shot pass@1 at 0.33 and 0.35, respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2406_02100
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Exploring Mathematical Extrapolation of Large Language Models with Synthetic Data
Li, Haolong
Ma, Yu
Zhang, Yinqi
Ye, Chen
Chen, Jie
Computation and Language
Large Language Models (LLMs) have shown excellent performance in language understanding, text generation, code synthesis, and many other tasks, while they still struggle in complex multi-step reasoning problems, such as mathematical reasoning. In this paper, through a newly proposed arithmetical puzzle problem, we show that the model can perform well on multi-step reasoning tasks via fine-tuning on high-quality synthetic data. Experimental results with the open-llama-3B model on three different test datasets show that not only the model can reach a zero-shot pass@1 at 0.44 on the in-domain dataset, it also demonstrates certain generalization capabilities on the out-of-domain datasets. Specifically, this paper has designed two out-of-domain datasets in the form of extending the numerical range and the composing components of the arithmetical puzzle problem separately. The fine-tuned models have shown encouraging performance on these two far more difficult tasks with the zero-shot pass@1 at 0.33 and 0.35, respectively.
title Exploring Mathematical Extrapolation of Large Language Models with Synthetic Data
topic Computation and Language
url https://arxiv.org/abs/2406.02100