Golden Touchstone: A Comprehensive Bilingual Benchmark for Evaluating Financial Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Xiaojun, Liu, Junxi, Su, Huanyi, Lin, Zhouchi, Qi, Yiyan, Xu, Chengjin, Su, Jiajun, Zhong, Jiajie, Wang, Fuwei, Wang, Saizhuo, Hua, Fengrui, Li, Jia, Guo, Jian
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911305562062848
author Wu, Xiaojun
Liu, Junxi
Su, Huanyi
Lin, Zhouchi
Qi, Yiyan
Xu, Chengjin
Su, Jiajun
Zhong, Jiajie
Wang, Fuwei
Wang, Saizhuo
Hua, Fengrui
Li, Jia
Guo, Jian
author_facet Wu, Xiaojun
Liu, Junxi
Su, Huanyi
Lin, Zhouchi
Qi, Yiyan
Xu, Chengjin
Su, Jiajun
Zhong, Jiajie
Wang, Fuwei
Wang, Saizhuo
Hua, Fengrui
Li, Jia
Guo, Jian
contents As large language models (LLMs) increasingly permeate the financial sector, there is a pressing need for a standardized method to comprehensively assess their performance. Existing financial benchmarks often suffer from limited language and task coverage, low-quality datasets, and inadequate adaptability for LLM evaluation. To address these limitations, we introduce Golden Touchstone, a comprehensive bilingual benchmark for financial LLMs, encompassing eight core financial NLP tasks in both Chinese and English. Developed from extensive open-source data collection and industry-specific demands, this benchmark thoroughly assesses models' language understanding and generation capabilities. Through comparative analysis of major models such as GPT-4o, Llama3, FinGPT, and FinMA, we reveal their strengths and limitations in processing complex financial information. Additionally, we open-source Touchstone-GPT, a financial LLM trained through continual pre-training and instruction tuning, which demonstrates strong performance on the bilingual benchmark but still has limitations in specific tasks. This research provides a practical evaluation tool for financial LLMs and guides future development and optimization. The source code for Golden Touchstone and model weight of Touchstone-GPT have been made publicly available at https://github.com/IDEA-FinAI/Golden-Touchstone.
format Preprint
id arxiv_https___arxiv_org_abs_2411_06272
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Golden Touchstone: A Comprehensive Bilingual Benchmark for Evaluating Financial Large Language Models
Wu, Xiaojun
Liu, Junxi
Su, Huanyi
Lin, Zhouchi
Qi, Yiyan
Xu, Chengjin
Su, Jiajun
Zhong, Jiajie
Wang, Fuwei
Wang, Saizhuo
Hua, Fengrui
Li, Jia
Guo, Jian
Computation and Language
Computational Engineering, Finance, and Science
As large language models (LLMs) increasingly permeate the financial sector, there is a pressing need for a standardized method to comprehensively assess their performance. Existing financial benchmarks often suffer from limited language and task coverage, low-quality datasets, and inadequate adaptability for LLM evaluation. To address these limitations, we introduce Golden Touchstone, a comprehensive bilingual benchmark for financial LLMs, encompassing eight core financial NLP tasks in both Chinese and English. Developed from extensive open-source data collection and industry-specific demands, this benchmark thoroughly assesses models' language understanding and generation capabilities. Through comparative analysis of major models such as GPT-4o, Llama3, FinGPT, and FinMA, we reveal their strengths and limitations in processing complex financial information. Additionally, we open-source Touchstone-GPT, a financial LLM trained through continual pre-training and instruction tuning, which demonstrates strong performance on the bilingual benchmark but still has limitations in specific tasks. This research provides a practical evaluation tool for financial LLMs and guides future development and optimization. The source code for Golden Touchstone and model weight of Touchstone-GPT have been made publicly available at https://github.com/IDEA-FinAI/Golden-Touchstone.
title Golden Touchstone: A Comprehensive Bilingual Benchmark for Evaluating Financial Large Language Models
topic Computation and Language
Computational Engineering, Finance, and Science
url https://arxiv.org/abs/2411.06272