FinMaster: A Holistic Benchmark for Mastering Full-Pipeline Financial Workflows with LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Junzhe, Yang, Chang, Cui, Aixin, Jin, Sihan, Wang, Ruiyu, Li, Bo, Huang, Xiao, Sun, Dongning, Wang, Xinrun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913848574869504
author Jiang, Junzhe
Yang, Chang
Cui, Aixin
Jin, Sihan
Wang, Ruiyu
Li, Bo
Huang, Xiao
Sun, Dongning
Wang, Xinrun
author_facet Jiang, Junzhe
Yang, Chang
Cui, Aixin
Jin, Sihan
Wang, Ruiyu
Li, Bo
Huang, Xiao
Sun, Dongning
Wang, Xinrun
contents Financial tasks are pivotal to global economic stability; however, their execution faces challenges including labor intensive processes, low error tolerance, data fragmentation, and tool limitations. Although large language models (LLMs) have succeeded in various natural language processing tasks and have shown potential in automating workflows through reasoning and contextual understanding, current benchmarks for evaluating LLMs in finance lack sufficient domain-specific data, have simplistic task design, and incomplete evaluation frameworks. To address these gaps, this article presents FinMaster, a comprehensive financial benchmark designed to systematically assess the capabilities of LLM in financial literacy, accounting, auditing, and consulting. Specifically, FinMaster comprises three main modules: i) FinSim, which builds simulators that generate synthetic, privacy-compliant financial data for companies to replicate market dynamics; ii) FinSuite, which provides tasks in core financial domains, spanning 183 tasks of various types and difficulty levels; and iii) FinEval, which develops a unified interface for evaluation. Extensive experiments over state-of-the-art LLMs reveal critical capability gaps in financial reasoning, with accuracy dropping from over 90% on basic tasks to merely 40% on complex scenarios requiring multi-step reasoning. This degradation exhibits the propagation of computational errors, where single-metric calculations initially demonstrating 58% accuracy decreased to 37% in multimetric scenarios. To the best of our knowledge, FinMaster is the first benchmark that covers full-pipeline financial workflows with challenging tasks. We hope that FinMaster can bridge the gap between research and industry practitioners, driving the adoption of LLMs in real-world financial practices to enhance efficiency and accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2505_13533
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FinMaster: A Holistic Benchmark for Mastering Full-Pipeline Financial Workflows with LLMs
Jiang, Junzhe
Yang, Chang
Cui, Aixin
Jin, Sihan
Wang, Ruiyu
Li, Bo
Huang, Xiao
Sun, Dongning
Wang, Xinrun
Artificial Intelligence
Machine Learning
General Finance
Financial tasks are pivotal to global economic stability; however, their execution faces challenges including labor intensive processes, low error tolerance, data fragmentation, and tool limitations. Although large language models (LLMs) have succeeded in various natural language processing tasks and have shown potential in automating workflows through reasoning and contextual understanding, current benchmarks for evaluating LLMs in finance lack sufficient domain-specific data, have simplistic task design, and incomplete evaluation frameworks. To address these gaps, this article presents FinMaster, a comprehensive financial benchmark designed to systematically assess the capabilities of LLM in financial literacy, accounting, auditing, and consulting. Specifically, FinMaster comprises three main modules: i) FinSim, which builds simulators that generate synthetic, privacy-compliant financial data for companies to replicate market dynamics; ii) FinSuite, which provides tasks in core financial domains, spanning 183 tasks of various types and difficulty levels; and iii) FinEval, which develops a unified interface for evaluation. Extensive experiments over state-of-the-art LLMs reveal critical capability gaps in financial reasoning, with accuracy dropping from over 90% on basic tasks to merely 40% on complex scenarios requiring multi-step reasoning. This degradation exhibits the propagation of computational errors, where single-metric calculations initially demonstrating 58% accuracy decreased to 37% in multimetric scenarios. To the best of our knowledge, FinMaster is the first benchmark that covers full-pipeline financial workflows with challenging tasks. We hope that FinMaster can bridge the gap between research and industry practitioners, driving the adoption of LLMs in real-world financial practices to enhance efficiency and accuracy.
title FinMaster: A Holistic Benchmark for Mastering Full-Pipeline Financial Workflows with LLMs
topic Artificial Intelligence
Machine Learning
General Finance
url https://arxiv.org/abs/2505.13533