FinAuditing: A Financial Taxonomy-Structured Multi-Document Benchmark for Evaluating LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911691310104576 |
|---|---|
| author | Wang, Yan Wang, Keyi Yang, Shanshan Patel, Jaisal Zhao, Jeff Mo, Fengran Peng, Xueqing Qian, Lingfei Chen, Yankai Gutiérrez-Basulto, Víctor Huang, Jimin Xiong, Guojun Liu, Xiao-Yang Liu, Xue Nie, Jian-Yun |
| author_facet | Wang, Yan Wang, Keyi Yang, Shanshan Patel, Jaisal Zhao, Jeff Mo, Fengran Peng, Xueqing Qian, Lingfei Chen, Yankai Gutiérrez-Basulto, Víctor Huang, Jimin Xiong, Guojun Liu, Xiao-Yang Liu, Xue Nie, Jian-Yun |
| contents | Going beyond simple text processing, financial auditing requires detecting semantic, structural, and numerical inconsistencies across large-scale disclosures. As financial reports are filed in XBRL, a structured XML format governed by accounting standards, auditing becomes a structured information extraction and reasoning problem involving concept alignment, taxonomy-defined relations, and cross-document consistency. Although large language models (LLMs) show promise on isolated financial tasks, their capability in professional-grade auditing remains unclear. We introduce FinAuditing, a taxonomy-aligned, structure-aware benchmark built from real XBRL filings. It contains 1,102 annotated instances averaging over 33k tokens and defines three tasks: Financial Semantic Matching (FinSM), Financial Relationship Extraction (FinRE), and Financial Mathematical Reasoning (FinMR). Evaluations of 13 state-of-the-art LLMs reveal substantial gaps in concept retrieval, taxonomy-aware relation modeling, and consistent cross-document reasoning. These findings highlight the need for realistic, structure-aware benchmarks. We release the evaluation code at https://github.com/The-FinAI/FinAuditing and the dataset at https://huggingface.co/collections/TheFinAI/finauditing. The task currently serves as the official benchmark of an ongoing public evaluation contest at https://open-finance-lab.github.io/SecureFinAI_Contest_2026/. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_08886 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | FinAuditing: A Financial Taxonomy-Structured Multi-Document Benchmark for Evaluating LLMs Wang, Yan Wang, Keyi Yang, Shanshan Patel, Jaisal Zhao, Jeff Mo, Fengran Peng, Xueqing Qian, Lingfei Chen, Yankai Gutiérrez-Basulto, Víctor Huang, Jimin Xiong, Guojun Liu, Xiao-Yang Liu, Xue Nie, Jian-Yun Computation and Language Computational Engineering, Finance, and Science Information Retrieval Going beyond simple text processing, financial auditing requires detecting semantic, structural, and numerical inconsistencies across large-scale disclosures. As financial reports are filed in XBRL, a structured XML format governed by accounting standards, auditing becomes a structured information extraction and reasoning problem involving concept alignment, taxonomy-defined relations, and cross-document consistency. Although large language models (LLMs) show promise on isolated financial tasks, their capability in professional-grade auditing remains unclear. We introduce FinAuditing, a taxonomy-aligned, structure-aware benchmark built from real XBRL filings. It contains 1,102 annotated instances averaging over 33k tokens and defines three tasks: Financial Semantic Matching (FinSM), Financial Relationship Extraction (FinRE), and Financial Mathematical Reasoning (FinMR). Evaluations of 13 state-of-the-art LLMs reveal substantial gaps in concept retrieval, taxonomy-aware relation modeling, and consistent cross-document reasoning. These findings highlight the need for realistic, structure-aware benchmarks. We release the evaluation code at https://github.com/The-FinAI/FinAuditing and the dataset at https://huggingface.co/collections/TheFinAI/finauditing. The task currently serves as the official benchmark of an ongoing public evaluation contest at https://open-finance-lab.github.io/SecureFinAI_Contest_2026/. |
| title | FinAuditing: A Financial Taxonomy-Structured Multi-Document Benchmark for Evaluating LLMs |
| topic | Computation and Language Computational Engineering, Finance, and Science Information Retrieval |
| url | https://arxiv.org/abs/2510.08886 |