FinAuditing: A Financial Taxonomy-Structured Multi-Document Benchmark for Evaluating LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Yan, Wang, Keyi, Yang, Shanshan, Patel, Jaisal, Zhao, Jeff, Mo, Fengran, Peng, Xueqing, Qian, Lingfei, Chen, Yankai, Gutiérrez-Basulto, Víctor, Huang, Jimin, Xiong, Guojun, Liu, Xiao-Yang, Liu, Xue, Nie, Jian-Yun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911691310104576
author Wang, Yan
Wang, Keyi
Yang, Shanshan
Patel, Jaisal
Zhao, Jeff
Mo, Fengran
Peng, Xueqing
Qian, Lingfei
Chen, Yankai
Gutiérrez-Basulto, Víctor
Huang, Jimin
Xiong, Guojun
Liu, Xiao-Yang
Liu, Xue
Nie, Jian-Yun
author_facet Wang, Yan
Wang, Keyi
Yang, Shanshan
Patel, Jaisal
Zhao, Jeff
Mo, Fengran
Peng, Xueqing
Qian, Lingfei
Chen, Yankai
Gutiérrez-Basulto, Víctor
Huang, Jimin
Xiong, Guojun
Liu, Xiao-Yang
Liu, Xue
Nie, Jian-Yun
contents Going beyond simple text processing, financial auditing requires detecting semantic, structural, and numerical inconsistencies across large-scale disclosures. As financial reports are filed in XBRL, a structured XML format governed by accounting standards, auditing becomes a structured information extraction and reasoning problem involving concept alignment, taxonomy-defined relations, and cross-document consistency. Although large language models (LLMs) show promise on isolated financial tasks, their capability in professional-grade auditing remains unclear. We introduce FinAuditing, a taxonomy-aligned, structure-aware benchmark built from real XBRL filings. It contains 1,102 annotated instances averaging over 33k tokens and defines three tasks: Financial Semantic Matching (FinSM), Financial Relationship Extraction (FinRE), and Financial Mathematical Reasoning (FinMR). Evaluations of 13 state-of-the-art LLMs reveal substantial gaps in concept retrieval, taxonomy-aware relation modeling, and consistent cross-document reasoning. These findings highlight the need for realistic, structure-aware benchmarks. We release the evaluation code at https://github.com/The-FinAI/FinAuditing and the dataset at https://huggingface.co/collections/TheFinAI/finauditing. The task currently serves as the official benchmark of an ongoing public evaluation contest at https://open-finance-lab.github.io/SecureFinAI_Contest_2026/.
format Preprint
id arxiv_https___arxiv_org_abs_2510_08886
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FinAuditing: A Financial Taxonomy-Structured Multi-Document Benchmark for Evaluating LLMs
Wang, Yan
Wang, Keyi
Yang, Shanshan
Patel, Jaisal
Zhao, Jeff
Mo, Fengran
Peng, Xueqing
Qian, Lingfei
Chen, Yankai
Gutiérrez-Basulto, Víctor
Huang, Jimin
Xiong, Guojun
Liu, Xiao-Yang
Liu, Xue
Nie, Jian-Yun
Computation and Language
Computational Engineering, Finance, and Science
Information Retrieval
Going beyond simple text processing, financial auditing requires detecting semantic, structural, and numerical inconsistencies across large-scale disclosures. As financial reports are filed in XBRL, a structured XML format governed by accounting standards, auditing becomes a structured information extraction and reasoning problem involving concept alignment, taxonomy-defined relations, and cross-document consistency. Although large language models (LLMs) show promise on isolated financial tasks, their capability in professional-grade auditing remains unclear. We introduce FinAuditing, a taxonomy-aligned, structure-aware benchmark built from real XBRL filings. It contains 1,102 annotated instances averaging over 33k tokens and defines three tasks: Financial Semantic Matching (FinSM), Financial Relationship Extraction (FinRE), and Financial Mathematical Reasoning (FinMR). Evaluations of 13 state-of-the-art LLMs reveal substantial gaps in concept retrieval, taxonomy-aware relation modeling, and consistent cross-document reasoning. These findings highlight the need for realistic, structure-aware benchmarks. We release the evaluation code at https://github.com/The-FinAI/FinAuditing and the dataset at https://huggingface.co/collections/TheFinAI/finauditing. The task currently serves as the official benchmark of an ongoing public evaluation contest at https://open-finance-lab.github.io/SecureFinAI_Contest_2026/.
title FinAuditing: A Financial Taxonomy-Structured Multi-Document Benchmark for Evaluating LLMs
topic Computation and Language
Computational Engineering, Finance, and Science
Information Retrieval
url https://arxiv.org/abs/2510.08886