Salvato in:
Dettagli Bibliografici
Autori principali: Gupta, Jatin, Sharma, Akhil, Singhania, Saransh, Adnan, Mohammad, Deo, Sakshi, Abidi, Ali Imam, Gupta, Keshav
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:https://arxiv.org/abs/2506.21031
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909660584345600
author Gupta, Jatin
Sharma, Akhil
Singhania, Saransh
Adnan, Mohammad
Deo, Sakshi
Abidi, Ali Imam
Gupta, Keshav
author_facet Gupta, Jatin
Sharma, Akhil
Singhania, Saransh
Adnan, Mohammad
Deo, Sakshi
Abidi, Ali Imam
Gupta, Keshav
contents Advanced intelligent systems, particularly Large Language Models (LLMs), are significantly reshaping financial practices through advancements in Natural Language Processing (NLP). However, the extent to which these models effectively capture and apply domain-specific financial knowledge remains uncertain. Addressing a critical gap in the expansive Indian financial context, this paper introduces CA-Ben, a Chartered Accountancy benchmark specifically designed to evaluate the financial, legal, and quantitative reasoning capabilities of LLMs. CA-Ben comprises structured question-answer datasets derived from the rigorous examinations conducted by the Institute of Chartered Accountants of India (ICAI), spanning foundational, intermediate, and advanced CA curriculum stages. Six prominent LLMs i.e. GPT 4o, LLAMA 3.3 70B, LLAMA 3.1 405B, MISTRAL Large, Claude 3.5 Sonnet, and Microsoft Phi 4 were evaluated using standardized protocols. Results indicate variations in performance, with Claude 3.5 Sonnet and GPT-4o outperforming others, especially in conceptual and legal reasoning. Notable challenges emerged in numerical computations and legal interpretations. The findings emphasize the strengths and limitations of current LLMs, suggesting future improvements through hybrid reasoning and retrieval-augmented generation methods, particularly for quantitative analysis and accurate legal interpretation.
format Preprint
id arxiv_https___arxiv_org_abs_2506_21031
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Large Language Models Acing Chartered Accountancy
Gupta, Jatin
Sharma, Akhil
Singhania, Saransh
Adnan, Mohammad
Deo, Sakshi
Abidi, Ali Imam
Gupta, Keshav
Computation and Language
Artificial Intelligence
Advanced intelligent systems, particularly Large Language Models (LLMs), are significantly reshaping financial practices through advancements in Natural Language Processing (NLP). However, the extent to which these models effectively capture and apply domain-specific financial knowledge remains uncertain. Addressing a critical gap in the expansive Indian financial context, this paper introduces CA-Ben, a Chartered Accountancy benchmark specifically designed to evaluate the financial, legal, and quantitative reasoning capabilities of LLMs. CA-Ben comprises structured question-answer datasets derived from the rigorous examinations conducted by the Institute of Chartered Accountants of India (ICAI), spanning foundational, intermediate, and advanced CA curriculum stages. Six prominent LLMs i.e. GPT 4o, LLAMA 3.3 70B, LLAMA 3.1 405B, MISTRAL Large, Claude 3.5 Sonnet, and Microsoft Phi 4 were evaluated using standardized protocols. Results indicate variations in performance, with Claude 3.5 Sonnet and GPT-4o outperforming others, especially in conceptual and legal reasoning. Notable challenges emerged in numerical computations and legal interpretations. The findings emphasize the strengths and limitations of current LLMs, suggesting future improvements through hybrid reasoning and retrieval-augmented generation methods, particularly for quantitative analysis and accurate legal interpretation.
title Large Language Models Acing Chartered Accountancy
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2506.21031