Pre-Calc: Learning to Use the Calculator Improves Numeracy in Language Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Veerendranath, Vishruth, Shah, Vishwa, Ghate, Kshitish
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914849038008320
author Veerendranath, Vishruth
Shah, Vishwa
Ghate, Kshitish
author_facet Veerendranath, Vishruth
Shah, Vishwa
Ghate, Kshitish
contents Quantitative and numerical comprehension in language is an important task in many fields like education and finance, but still remains a challenging task for language models. While tool and calculator usage has shown to be helpful to improve mathematical reasoning in large pretrained decoder-only language models, this remains unexplored for smaller language models with encoders. In this paper, we propose Pre-Calc, a simple pre-finetuning objective of learning to use the calculator for both encoder-only and encoder-decoder architectures, formulated as a discriminative and generative task respectively. We pre-train BERT and RoBERTa for discriminative calculator use and Flan-T5 for generative calculator use on the MAWPS, SVAMP, and AsDiv-A datasets, which improves performance on downstream tasks that require numerical understanding. Our code and data are available at https://github.com/calc-cmu/pre-calc.
format Preprint
id arxiv_https___arxiv_org_abs_2404_14355
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Pre-Calc: Learning to Use the Calculator Improves Numeracy in Language Models
Veerendranath, Vishruth
Shah, Vishwa
Ghate, Kshitish
Computation and Language
Artificial Intelligence
Quantitative and numerical comprehension in language is an important task in many fields like education and finance, but still remains a challenging task for language models. While tool and calculator usage has shown to be helpful to improve mathematical reasoning in large pretrained decoder-only language models, this remains unexplored for smaller language models with encoders. In this paper, we propose Pre-Calc, a simple pre-finetuning objective of learning to use the calculator for both encoder-only and encoder-decoder architectures, formulated as a discriminative and generative task respectively. We pre-train BERT and RoBERTa for discriminative calculator use and Flan-T5 for generative calculator use on the MAWPS, SVAMP, and AsDiv-A datasets, which improves performance on downstream tasks that require numerical understanding. Our code and data are available at https://github.com/calc-cmu/pre-calc.
title Pre-Calc: Learning to Use the Calculator Improves Numeracy in Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2404.14355