Towards Automated Regulatory Compliance Verification in Financial Auditing with Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866915405050675200 |
|---|---|
| author | Berger, Armin Hillebrand, Lars Leonhard, David Deußer, Tobias de Oliveira, Thiago Bell Felix Dilmaghani, Tim Khaled, Mohamed Kliem, Bernd Loitz, Rüdiger Bauckhage, Christian Sifa, Rafet |
| author_facet | Berger, Armin Hillebrand, Lars Leonhard, David Deußer, Tobias de Oliveira, Thiago Bell Felix Dilmaghani, Tim Khaled, Mohamed Kliem, Bernd Loitz, Rüdiger Bauckhage, Christian Sifa, Rafet |
| contents | The auditing of financial documents, historically a labor-intensive process, stands on the precipice of transformation. AI-driven solutions have made inroads into streamlining this process by recommending pertinent text passages from financial reports to align with the legal requirements of accounting standards. However, a glaring limitation remains: these systems commonly fall short in verifying if the recommended excerpts indeed comply with the specific legal mandates. Hence, in this paper, we probe the efficiency of publicly available Large Language Models (LLMs) in the realm of regulatory compliance across different model configurations. We place particular emphasis on comparing cutting-edge open-source LLMs, such as Llama-2, with their proprietary counterparts like OpenAI's GPT models. This comparative analysis leverages two custom datasets provided by our partner PricewaterhouseCoopers (PwC) Germany. We find that the open-source Llama-2 70 billion model demonstrates outstanding performance in detecting non-compliance or true negative occurrences, beating all their proprietary counterparts. Nevertheless, proprietary models such as GPT-4 perform the best in a broad variety of scenarios, particularly in non-English contexts. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2507_16642 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Towards Automated Regulatory Compliance Verification in Financial Auditing with Large Language Models Berger, Armin Hillebrand, Lars Leonhard, David Deußer, Tobias de Oliveira, Thiago Bell Felix Dilmaghani, Tim Khaled, Mohamed Kliem, Bernd Loitz, Rüdiger Bauckhage, Christian Sifa, Rafet Computation and Language Artificial Intelligence Machine Learning The auditing of financial documents, historically a labor-intensive process, stands on the precipice of transformation. AI-driven solutions have made inroads into streamlining this process by recommending pertinent text passages from financial reports to align with the legal requirements of accounting standards. However, a glaring limitation remains: these systems commonly fall short in verifying if the recommended excerpts indeed comply with the specific legal mandates. Hence, in this paper, we probe the efficiency of publicly available Large Language Models (LLMs) in the realm of regulatory compliance across different model configurations. We place particular emphasis on comparing cutting-edge open-source LLMs, such as Llama-2, with their proprietary counterparts like OpenAI's GPT models. This comparative analysis leverages two custom datasets provided by our partner PricewaterhouseCoopers (PwC) Germany. We find that the open-source Llama-2 70 billion model demonstrates outstanding performance in detecting non-compliance or true negative occurrences, beating all their proprietary counterparts. Nevertheless, proprietary models such as GPT-4 perform the best in a broad variety of scenarios, particularly in non-English contexts. |
| title | Towards Automated Regulatory Compliance Verification in Financial Auditing with Large Language Models |
| topic | Computation and Language Artificial Intelligence Machine Learning |
| url | https://arxiv.org/abs/2507.16642 |