The ORCA Benchmark: Evaluating Real-World Calculation Accuracy in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Herambourg, Claudia, Siuda, Dawid, Kopczyńska, Julia, Santos, Joao R. L., Sas, Wojciech, Śmietańska-Nowak, Joanna
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912688156704768
author Herambourg, Claudia
Siuda, Dawid
Kopczyńska, Julia
Santos, Joao R. L.
Sas, Wojciech
Śmietańska-Nowak, Joanna
author_facet Herambourg, Claudia
Siuda, Dawid
Kopczyńska, Julia
Santos, Joao R. L.
Sas, Wojciech
Śmietańska-Nowak, Joanna
contents We present ORCA (Omni Research on Calculation in AI) Benchmark - a novel benchmark that evaluates large language models (LLMs) on multi-domain, real-life quantitative reasoning using verified outputs from Omni's calculator engine. In 500 natural-language tasks across domains such as finance, physics, health, and statistics, the five state-of-the-art systems (ChatGPT-5, Gemini~2.5~Flash, Claude~Sonnet~4.5, Grok~4, and DeepSeek~V3.2) achieved only $45\text{--}63\,\%$ accuracy, with errors mainly related to rounding ($35\,\%$) and calculation mistakes ($33\,\%$). Results in specific domains indicate strengths in mathematics and engineering, but weaknesses in physics and natural sciences. Correlation analysis ($r \approx 0.40\text{--}0.65$) shows that the models often fail together but differ in the types of errors they make, highlighting their partial complementarity rather than redundancy. Unlike standard math datasets, ORCA evaluates step-by-step reasoning, numerical precision, and domain generalization across real problems from finance, physics, health, and statistics.
format Preprint
id arxiv_https___arxiv_org_abs_2511_02589
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The ORCA Benchmark: Evaluating Real-World Calculation Accuracy in Large Language Models
Herambourg, Claudia
Siuda, Dawid
Kopczyńska, Julia
Santos, Joao R. L.
Sas, Wojciech
Śmietańska-Nowak, Joanna
Artificial Intelligence
We present ORCA (Omni Research on Calculation in AI) Benchmark - a novel benchmark that evaluates large language models (LLMs) on multi-domain, real-life quantitative reasoning using verified outputs from Omni's calculator engine. In 500 natural-language tasks across domains such as finance, physics, health, and statistics, the five state-of-the-art systems (ChatGPT-5, Gemini~2.5~Flash, Claude~Sonnet~4.5, Grok~4, and DeepSeek~V3.2) achieved only $45\text{--}63\,\%$ accuracy, with errors mainly related to rounding ($35\,\%$) and calculation mistakes ($33\,\%$). Results in specific domains indicate strengths in mathematics and engineering, but weaknesses in physics and natural sciences. Correlation analysis ($r \approx 0.40\text{--}0.65$) shows that the models often fail together but differ in the types of errors they make, highlighting their partial complementarity rather than redundancy. Unlike standard math datasets, ORCA evaluates step-by-step reasoning, numerical precision, and domain generalization across real problems from finance, physics, health, and statistics.
title The ORCA Benchmark: Evaluating Real-World Calculation Accuracy in Large Language Models
topic Artificial Intelligence
url https://arxiv.org/abs/2511.02589