EDINET-Bench: Evaluating LLMs on Complex Financial Tasks using Japanese Financial Statements

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sugiura, Issa, Ishida, Takashi, Makino, Taro, Tazuke, Chieko, Nakagawa, Takanori, Nakago, Kosuke, Ha, David
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911487263506432
author Sugiura, Issa
Ishida, Takashi
Makino, Taro
Tazuke, Chieko
Nakagawa, Takanori
Nakago, Kosuke
Ha, David
author_facet Sugiura, Issa
Ishida, Takashi
Makino, Taro
Tazuke, Chieko
Nakagawa, Takanori
Nakago, Kosuke
Ha, David
contents Large Language Models (LLMs) have made remarkable progress, surpassing human performance on several benchmarks in domains such as mathematics and coding. A key driver of this progress has been the development of benchmark datasets. In contrast, the financial domain poses higher entry barriers due to its demand for specialized expertise, and benchmarks remain relatively scarce compared to those in mathematics or coding. We introduce EDINET-Bench, an open-source Japanese financial benchmark designed to evaluate LLMs on challenging tasks such as accounting fraud detection, earnings forecasting, and industry classification. EDINET-Bench is constructed from ten years of annual reports filed by Japanese companies. These tasks require models to process entire annual reports and integrate information across multiple tables and textual sections, demanding expert-level reasoning that is challenging even for human professionals. Our experiments show that even state-of-the-art LLMs struggle in this domain, performing only marginally better than logistic regression in binary classification tasks such as fraud detection and earnings forecasting. Our results show that simply providing reports to LLMs in a straightforward setting is not enough. This highlights the need for benchmark frameworks that better reflect the environments in which financial professionals operate, with richer scaffolding such as realistic simulations and task-specific reasoning support to enable more effective problem solving. We make our dataset and code publicly available to support future research.
format Preprint
id arxiv_https___arxiv_org_abs_2506_08762
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EDINET-Bench: Evaluating LLMs on Complex Financial Tasks using Japanese Financial Statements
Sugiura, Issa
Ishida, Takashi
Makino, Taro
Tazuke, Chieko
Nakagawa, Takanori
Nakago, Kosuke
Ha, David
Statistical Finance
Computational Engineering, Finance, and Science
Computation and Language
Machine Learning
Large Language Models (LLMs) have made remarkable progress, surpassing human performance on several benchmarks in domains such as mathematics and coding. A key driver of this progress has been the development of benchmark datasets. In contrast, the financial domain poses higher entry barriers due to its demand for specialized expertise, and benchmarks remain relatively scarce compared to those in mathematics or coding. We introduce EDINET-Bench, an open-source Japanese financial benchmark designed to evaluate LLMs on challenging tasks such as accounting fraud detection, earnings forecasting, and industry classification. EDINET-Bench is constructed from ten years of annual reports filed by Japanese companies. These tasks require models to process entire annual reports and integrate information across multiple tables and textual sections, demanding expert-level reasoning that is challenging even for human professionals. Our experiments show that even state-of-the-art LLMs struggle in this domain, performing only marginally better than logistic regression in binary classification tasks such as fraud detection and earnings forecasting. Our results show that simply providing reports to LLMs in a straightforward setting is not enough. This highlights the need for benchmark frameworks that better reflect the environments in which financial professionals operate, with richer scaffolding such as realistic simulations and task-specific reasoning support to enable more effective problem solving. We make our dataset and code publicly available to support future research.
title EDINET-Bench: Evaluating LLMs on Complex Financial Tasks using Japanese Financial Statements
topic Statistical Finance
Computational Engineering, Finance, and Science
Computation and Language
Machine Learning
url https://arxiv.org/abs/2506.08762