FinGAIA: A Chinese Benchmark for AI Agents in Real-World Financial Domain

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zeng, Lingfeng, Lou, Fangqi, Wang, Zixuan, Xu, Jiajie, Niu, Jinyi, Li, Mengping, Dong, Yifan, Qi, Qi, Zhang, Wei, Yang, Ziwei, Han, Jun, Feng, Ruilun, Hu, Ruiqi, Zhang, Lejie, Feng, Zhengbo, Ren, Yicheng, Guo, Xin, Liu, Zhaowei, Cheng, Dongpo, Cai, Weige, Zhang, Liwen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913967002091520
author Zeng, Lingfeng
Lou, Fangqi
Wang, Zixuan
Xu, Jiajie
Niu, Jinyi
Li, Mengping
Dong, Yifan
Qi, Qi
Zhang, Wei
Yang, Ziwei
Han, Jun
Feng, Ruilun
Hu, Ruiqi
Zhang, Lejie
Feng, Zhengbo
Ren, Yicheng
Guo, Xin
Liu, Zhaowei
Cheng, Dongpo
Cai, Weige
Zhang, Liwen
author_facet Zeng, Lingfeng
Lou, Fangqi
Wang, Zixuan
Xu, Jiajie
Niu, Jinyi
Li, Mengping
Dong, Yifan
Qi, Qi
Zhang, Wei
Yang, Ziwei
Han, Jun
Feng, Ruilun
Hu, Ruiqi
Zhang, Lejie
Feng, Zhengbo
Ren, Yicheng
Guo, Xin
Liu, Zhaowei
Cheng, Dongpo
Cai, Weige
Zhang, Liwen
contents The booming development of AI agents presents unprecedented opportunities for automating complex tasks across various domains. However, their multi-step, multi-tool collaboration capabilities in the financial sector remain underexplored. This paper introduces FinGAIA, an end-to-end benchmark designed to evaluate the practical abilities of AI agents in the financial domain. FinGAIA comprises 407 meticulously crafted tasks, spanning seven major financial sub-domains: securities, funds, banking, insurance, futures, trusts, and asset management. These tasks are organized into three hierarchical levels of scenario depth: basic business analysis, asset decision support, and strategic risk management. We evaluated 10 mainstream AI agents in a zero-shot setting. The best-performing agent, ChatGPT, achieved an overall accuracy of 48.9\%, which, while superior to non-professionals, still lags financial experts by over 35 percentage points. Error analysis has revealed five recurring failure patterns: Cross-modal Alignment Deficiency, Financial Terminological Bias, Operational Process Awareness Barrier, among others. These patterns point to crucial directions for future research. Our work provides the first agent benchmark closely related to the financial domain, aiming to objectively assess and promote the development of agents in this crucial field. Partial data is available at https://github.com/SUFE-AIFLM-Lab/FinGAIA.
format Preprint
id arxiv_https___arxiv_org_abs_2507_17186
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FinGAIA: A Chinese Benchmark for AI Agents in Real-World Financial Domain
Zeng, Lingfeng
Lou, Fangqi
Wang, Zixuan
Xu, Jiajie
Niu, Jinyi
Li, Mengping
Dong, Yifan
Qi, Qi
Zhang, Wei
Yang, Ziwei
Han, Jun
Feng, Ruilun
Hu, Ruiqi
Zhang, Lejie
Feng, Zhengbo
Ren, Yicheng
Guo, Xin
Liu, Zhaowei
Cheng, Dongpo
Cai, Weige
Zhang, Liwen
Computation and Language
The booming development of AI agents presents unprecedented opportunities for automating complex tasks across various domains. However, their multi-step, multi-tool collaboration capabilities in the financial sector remain underexplored. This paper introduces FinGAIA, an end-to-end benchmark designed to evaluate the practical abilities of AI agents in the financial domain. FinGAIA comprises 407 meticulously crafted tasks, spanning seven major financial sub-domains: securities, funds, banking, insurance, futures, trusts, and asset management. These tasks are organized into three hierarchical levels of scenario depth: basic business analysis, asset decision support, and strategic risk management. We evaluated 10 mainstream AI agents in a zero-shot setting. The best-performing agent, ChatGPT, achieved an overall accuracy of 48.9\%, which, while superior to non-professionals, still lags financial experts by over 35 percentage points. Error analysis has revealed five recurring failure patterns: Cross-modal Alignment Deficiency, Financial Terminological Bias, Operational Process Awareness Barrier, among others. These patterns point to crucial directions for future research. Our work provides the first agent benchmark closely related to the financial domain, aiming to objectively assess and promote the development of agents in this crucial field. Partial data is available at https://github.com/SUFE-AIFLM-Lab/FinGAIA.
title FinGAIA: A Chinese Benchmark for AI Agents in Real-World Financial Domain
topic Computation and Language
url https://arxiv.org/abs/2507.17186