How Numerical Precision Affects Arithmetical Reasoning Capabilities of LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Feng, Guhao, Yang, Kai, Gu, Yuntian, Ai, Xinyue, Luo, Shengjie, Sun, Jiacheng, He, Di, Li, Zhenguo, Wang, Liwei
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918067054837760
author Feng, Guhao
Yang, Kai
Gu, Yuntian
Ai, Xinyue
Luo, Shengjie
Sun, Jiacheng
He, Di
Li, Zhenguo
Wang, Liwei
author_facet Feng, Guhao
Yang, Kai
Gu, Yuntian
Ai, Xinyue
Luo, Shengjie
Sun, Jiacheng
He, Di
Li, Zhenguo
Wang, Liwei
contents Despite the remarkable success of Transformer-based large language models (LLMs) across various domains, understanding and enhancing their mathematical capabilities remains a significant challenge. In this paper, we conduct a rigorous theoretical analysis of LLMs' mathematical abilities, with a specific focus on their arithmetic performances. We identify numerical precision as a key factor that influences their effectiveness in arithmetical tasks. Our results show that Transformers operating with low numerical precision fail to address arithmetic tasks, such as iterated addition and integer multiplication, unless the model size grows super-polynomially with respect to the input length. In contrast, Transformers with standard numerical precision can efficiently handle these tasks with significantly smaller model sizes. We further support our theoretical findings through empirical experiments that explore the impact of varying numerical precision on arithmetic tasks, providing valuable insights for improving the mathematical reasoning capabilities of LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2410_13857
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle How Numerical Precision Affects Arithmetical Reasoning Capabilities of LLMs
Feng, Guhao
Yang, Kai
Gu, Yuntian
Ai, Xinyue
Luo, Shengjie
Sun, Jiacheng
He, Di
Li, Zhenguo
Wang, Liwei
Machine Learning
Artificial Intelligence
Computation and Language
Despite the remarkable success of Transformer-based large language models (LLMs) across various domains, understanding and enhancing their mathematical capabilities remains a significant challenge. In this paper, we conduct a rigorous theoretical analysis of LLMs' mathematical abilities, with a specific focus on their arithmetic performances. We identify numerical precision as a key factor that influences their effectiveness in arithmetical tasks. Our results show that Transformers operating with low numerical precision fail to address arithmetic tasks, such as iterated addition and integer multiplication, unless the model size grows super-polynomially with respect to the input length. In contrast, Transformers with standard numerical precision can efficiently handle these tasks with significantly smaller model sizes. We further support our theoretical findings through empirical experiments that explore the impact of varying numerical precision on arithmetic tasks, providing valuable insights for improving the mathematical reasoning capabilities of LLMs.
title How Numerical Precision Affects Arithmetical Reasoning Capabilities of LLMs
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2410.13857