Can LLMs $\textit{understand}$ Math? -- Exploring the Pitfalls in Mathematical Reasoning
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866909618789154816 |
|---|---|
| author | Roy, Tiasa Singha Baral, Aditeya Jhaveri, Ayush Rajesh Baig, Yusuf |
| author_facet | Roy, Tiasa Singha Baral, Aditeya Jhaveri, Ayush Rajesh Baig, Yusuf |
| contents | Large language models (LLMs) demonstrate considerable potential in various natural language tasks but face significant challenges in mathematical reasoning, particularly in executing precise, multi-step logic. However, current evaluation frameworks judge their performance solely based on accuracy, which only accounts for the final answer. This study explores these pitfalls by employing a novel evaluation framework. We propose an evaluation metric called the MAPLE score, which holistically quantifies reasoning misalignment by integrating error rates, redundancy, and validity. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_15623 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Can LLMs $\textit{understand}$ Math? -- Exploring the Pitfalls in Mathematical Reasoning Roy, Tiasa Singha Baral, Aditeya Jhaveri, Ayush Rajesh Baig, Yusuf Computation and Language Machine Learning Large language models (LLMs) demonstrate considerable potential in various natural language tasks but face significant challenges in mathematical reasoning, particularly in executing precise, multi-step logic. However, current evaluation frameworks judge their performance solely based on accuracy, which only accounts for the final answer. This study explores these pitfalls by employing a novel evaluation framework. We propose an evaluation metric called the MAPLE score, which holistically quantifies reasoning misalignment by integrating error rates, redundancy, and validity. |
| title | Can LLMs $\textit{understand}$ Math? -- Exploring the Pitfalls in Mathematical Reasoning |
| topic | Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2505.15623 |