Can LLMs $\textit{understand}$ Math? -- Exploring the Pitfalls in Mathematical Reasoning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Roy, Tiasa Singha, Baral, Aditeya, Jhaveri, Ayush Rajesh, Baig, Yusuf
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909618789154816
author Roy, Tiasa Singha
Baral, Aditeya
Jhaveri, Ayush Rajesh
Baig, Yusuf
author_facet Roy, Tiasa Singha
Baral, Aditeya
Jhaveri, Ayush Rajesh
Baig, Yusuf
contents Large language models (LLMs) demonstrate considerable potential in various natural language tasks but face significant challenges in mathematical reasoning, particularly in executing precise, multi-step logic. However, current evaluation frameworks judge their performance solely based on accuracy, which only accounts for the final answer. This study explores these pitfalls by employing a novel evaluation framework. We propose an evaluation metric called the MAPLE score, which holistically quantifies reasoning misalignment by integrating error rates, redundancy, and validity.
format Preprint
id arxiv_https___arxiv_org_abs_2505_15623
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Can LLMs $\textit{understand}$ Math? -- Exploring the Pitfalls in Mathematical Reasoning
Roy, Tiasa Singha
Baral, Aditeya
Jhaveri, Ayush Rajesh
Baig, Yusuf
Computation and Language
Machine Learning
Large language models (LLMs) demonstrate considerable potential in various natural language tasks but face significant challenges in mathematical reasoning, particularly in executing precise, multi-step logic. However, current evaluation frameworks judge their performance solely based on accuracy, which only accounts for the final answer. This study explores these pitfalls by employing a novel evaluation framework. We propose an evaluation metric called the MAPLE score, which holistically quantifies reasoning misalignment by integrating error rates, redundancy, and validity.
title Can LLMs $\textit{understand}$ Math? -- Exploring the Pitfalls in Mathematical Reasoning
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2505.15623