FinEval: A Chinese Financial Domain Knowledge Evaluation Benchmark for Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guo, Xin, Xia, Haotian, Liu, Zhaowei, Cao, Hanyang, Yang, Zhi, Liu, Zhiqiang, Wang, Sizhe, Niu, Jinyi, Wang, Chuqi, Wang, Yanhui, Liang, Xiaolong, Huang, Xiaoming, Zhu, Bing, Wei, Zhongyu, Chen, Yun, Shen, Weining, Zhang, Liwen
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912147403964416
author Guo, Xin
Xia, Haotian
Liu, Zhaowei
Cao, Hanyang
Yang, Zhi
Liu, Zhiqiang
Wang, Sizhe
Niu, Jinyi
Wang, Chuqi
Wang, Yanhui
Liang, Xiaolong
Huang, Xiaoming
Zhu, Bing
Wei, Zhongyu
Chen, Yun
Shen, Weining
Zhang, Liwen
author_facet Guo, Xin
Xia, Haotian
Liu, Zhaowei
Cao, Hanyang
Yang, Zhi
Liu, Zhiqiang
Wang, Sizhe
Niu, Jinyi
Wang, Chuqi
Wang, Yanhui
Liang, Xiaolong
Huang, Xiaoming
Zhu, Bing
Wei, Zhongyu
Chen, Yun
Shen, Weining
Zhang, Liwen
contents Large language models have demonstrated outstanding performance in various natural language processing tasks, but their security capabilities in the financial domain have not been explored, and their performance on complex tasks like financial agent remains unknown. This paper presents FinEval, a benchmark designed to evaluate LLMs' financial domain knowledge and practical abilities. The dataset contains 8,351 questions categorized into four different key areas: Financial Academic Knowledge, Financial Industry Knowledge, Financial Security Knowledge, and Financial Agent. Financial Academic Knowledge comprises 4,661 multiple-choice questions spanning 34 subjects such as finance and economics. Financial Industry Knowledge contains 1,434 questions covering practical scenarios like investment research. Financial Security Knowledge assesses models through 1,640 questions on topics like application security and cryptography. Financial Agent evaluates tool usage and complex reasoning with 616 questions. FinEval has multiple evaluation settings, including zero-shot, five-shot with chain-of-thought, and assesses model performance using objective and subjective criteria. Our results show that Claude 3.5-Sonnet achieves the highest weighted average score of 72.9 across all financial domain categories under zero-shot setting. Our work provides a comprehensive benchmark closely aligned with Chinese financial domain.
format Preprint
id arxiv_https___arxiv_org_abs_2308_09975
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle FinEval: A Chinese Financial Domain Knowledge Evaluation Benchmark for Large Language Models
Guo, Xin
Xia, Haotian
Liu, Zhaowei
Cao, Hanyang
Yang, Zhi
Liu, Zhiqiang
Wang, Sizhe
Niu, Jinyi
Wang, Chuqi
Wang, Yanhui
Liang, Xiaolong
Huang, Xiaoming
Zhu, Bing
Wei, Zhongyu
Chen, Yun
Shen, Weining
Zhang, Liwen
Computation and Language
Large language models have demonstrated outstanding performance in various natural language processing tasks, but their security capabilities in the financial domain have not been explored, and their performance on complex tasks like financial agent remains unknown. This paper presents FinEval, a benchmark designed to evaluate LLMs' financial domain knowledge and practical abilities. The dataset contains 8,351 questions categorized into four different key areas: Financial Academic Knowledge, Financial Industry Knowledge, Financial Security Knowledge, and Financial Agent. Financial Academic Knowledge comprises 4,661 multiple-choice questions spanning 34 subjects such as finance and economics. Financial Industry Knowledge contains 1,434 questions covering practical scenarios like investment research. Financial Security Knowledge assesses models through 1,640 questions on topics like application security and cryptography. Financial Agent evaluates tool usage and complex reasoning with 616 questions. FinEval has multiple evaluation settings, including zero-shot, five-shot with chain-of-thought, and assesses model performance using objective and subjective criteria. Our results show that Claude 3.5-Sonnet achieves the highest weighted average score of 72.9 across all financial domain categories under zero-shot setting. Our work provides a comprehensive benchmark closely aligned with Chinese financial domain.
title FinEval: A Chinese Financial Domain Knowledge Evaluation Benchmark for Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2308.09975