Can LLMs subtract numbers?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jobanputra, Mayank, Walter, Nils Philipp, Mehta, Maitrey, Veseli, Blerta, Chapple, Evan Parker Kelly, Wang, Yifan, Chetani, Sneha, Pavlick, Ellie, Vergari, Antonio, Demberg, Vera
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912687920775168
author Jobanputra, Mayank
Walter, Nils Philipp
Mehta, Maitrey
Veseli, Blerta
Chapple, Evan Parker Kelly
Wang, Yifan
Chetani, Sneha
Pavlick, Ellie
Vergari, Antonio
Demberg, Vera
author_facet Jobanputra, Mayank
Walter, Nils Philipp
Mehta, Maitrey
Veseli, Blerta
Chapple, Evan Parker Kelly
Wang, Yifan
Chetani, Sneha
Pavlick, Ellie
Vergari, Antonio
Demberg, Vera
contents We present a systematic study of subtraction in large language models (LLMs). While prior benchmarks emphasize addition and multiplication, subtraction has received comparatively little attention despite being structurally distinct as a non-commutative operation. We evaluate eight pretrained LLMs spanning four families on addition and subtraction problems. Our experiments reveal that subtraction accuracy lags behind addition by a wide margin. We find that the errors for ($a-b$) are concentrated in cases where ($a<b$). In such cases, LLMs frequently produce the correct magnitude but omit the negative sign. Probing analyses show that LLMs internally encode whether results should be negative, yet this information is often not reflected in generated outputs. We further test well-known techniques such as few-shot learning and instruction-tuning to see if they can improve the LLMs' performance. Our results suggest that while few-shot prompting yields modest gains, the instruction-tuned models achieve near-perfect accuracies in generating the negative sign. Together, these findings provide a clearer characterization of the limitations and recoverability of LLMs' arithmetic capabilities in subtraction.
format Preprint
id arxiv_https___arxiv_org_abs_2511_02795
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Can LLMs subtract numbers?
Jobanputra, Mayank
Walter, Nils Philipp
Mehta, Maitrey
Veseli, Blerta
Chapple, Evan Parker Kelly
Wang, Yifan
Chetani, Sneha
Pavlick, Ellie
Vergari, Antonio
Demberg, Vera
Machine Learning
Computation and Language
We present a systematic study of subtraction in large language models (LLMs). While prior benchmarks emphasize addition and multiplication, subtraction has received comparatively little attention despite being structurally distinct as a non-commutative operation. We evaluate eight pretrained LLMs spanning four families on addition and subtraction problems. Our experiments reveal that subtraction accuracy lags behind addition by a wide margin. We find that the errors for ($a-b$) are concentrated in cases where ($a<b$). In such cases, LLMs frequently produce the correct magnitude but omit the negative sign. Probing analyses show that LLMs internally encode whether results should be negative, yet this information is often not reflected in generated outputs. We further test well-known techniques such as few-shot learning and instruction-tuning to see if they can improve the LLMs' performance. Our results suggest that while few-shot prompting yields modest gains, the instruction-tuned models achieve near-perfect accuracies in generating the negative sign. Together, these findings provide a clearer characterization of the limitations and recoverability of LLMs' arithmetic capabilities in subtraction.
title Can LLMs subtract numbers?
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2511.02795