The ORCA Benchmark: Evaluating Real-World Calculation Accuracy in Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912688156704768 |
|---|---|
| author | Herambourg, Claudia Siuda, Dawid Kopczyńska, Julia Santos, Joao R. L. Sas, Wojciech Śmietańska-Nowak, Joanna |
| author_facet | Herambourg, Claudia Siuda, Dawid Kopczyńska, Julia Santos, Joao R. L. Sas, Wojciech Śmietańska-Nowak, Joanna |
| contents | We present ORCA (Omni Research on Calculation in AI) Benchmark - a novel benchmark that evaluates large language models (LLMs) on multi-domain, real-life quantitative reasoning using verified outputs from Omni's calculator engine. In 500 natural-language tasks across domains such as finance, physics, health, and statistics, the five state-of-the-art systems (ChatGPT-5, Gemini~2.5~Flash, Claude~Sonnet~4.5, Grok~4, and DeepSeek~V3.2) achieved only $45\text{--}63\,\%$ accuracy, with errors mainly related to rounding ($35\,\%$) and calculation mistakes ($33\,\%$). Results in specific domains indicate strengths in mathematics and engineering, but weaknesses in physics and natural sciences. Correlation analysis ($r \approx 0.40\text{--}0.65$) shows that the models often fail together but differ in the types of errors they make, highlighting their partial complementarity rather than redundancy. Unlike standard math datasets, ORCA evaluates step-by-step reasoning, numerical precision, and domain generalization across real problems from finance, physics, health, and statistics. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_02589 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | The ORCA Benchmark: Evaluating Real-World Calculation Accuracy in Large Language Models Herambourg, Claudia Siuda, Dawid Kopczyńska, Julia Santos, Joao R. L. Sas, Wojciech Śmietańska-Nowak, Joanna Artificial Intelligence We present ORCA (Omni Research on Calculation in AI) Benchmark - a novel benchmark that evaluates large language models (LLMs) on multi-domain, real-life quantitative reasoning using verified outputs from Omni's calculator engine. In 500 natural-language tasks across domains such as finance, physics, health, and statistics, the five state-of-the-art systems (ChatGPT-5, Gemini~2.5~Flash, Claude~Sonnet~4.5, Grok~4, and DeepSeek~V3.2) achieved only $45\text{--}63\,\%$ accuracy, with errors mainly related to rounding ($35\,\%$) and calculation mistakes ($33\,\%$). Results in specific domains indicate strengths in mathematics and engineering, but weaknesses in physics and natural sciences. Correlation analysis ($r \approx 0.40\text{--}0.65$) shows that the models often fail together but differ in the types of errors they make, highlighting their partial complementarity rather than redundancy. Unlike standard math datasets, ORCA evaluates step-by-step reasoning, numerical precision, and domain generalization across real problems from finance, physics, health, and statistics. |
| title | The ORCA Benchmark: Evaluating Real-World Calculation Accuracy in Large Language Models |
| topic | Artificial Intelligence |
| url | https://arxiv.org/abs/2511.02589 |