Golden Touchstone: A Comprehensive Bilingual Benchmark for Evaluating Financial Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911305562062848 |
|---|---|
| author | Wu, Xiaojun Liu, Junxi Su, Huanyi Lin, Zhouchi Qi, Yiyan Xu, Chengjin Su, Jiajun Zhong, Jiajie Wang, Fuwei Wang, Saizhuo Hua, Fengrui Li, Jia Guo, Jian |
| author_facet | Wu, Xiaojun Liu, Junxi Su, Huanyi Lin, Zhouchi Qi, Yiyan Xu, Chengjin Su, Jiajun Zhong, Jiajie Wang, Fuwei Wang, Saizhuo Hua, Fengrui Li, Jia Guo, Jian |
| contents | As large language models (LLMs) increasingly permeate the financial sector, there is a pressing need for a standardized method to comprehensively assess their performance. Existing financial benchmarks often suffer from limited language and task coverage, low-quality datasets, and inadequate adaptability for LLM evaluation. To address these limitations, we introduce Golden Touchstone, a comprehensive bilingual benchmark for financial LLMs, encompassing eight core financial NLP tasks in both Chinese and English. Developed from extensive open-source data collection and industry-specific demands, this benchmark thoroughly assesses models' language understanding and generation capabilities. Through comparative analysis of major models such as GPT-4o, Llama3, FinGPT, and FinMA, we reveal their strengths and limitations in processing complex financial information. Additionally, we open-source Touchstone-GPT, a financial LLM trained through continual pre-training and instruction tuning, which demonstrates strong performance on the bilingual benchmark but still has limitations in specific tasks. This research provides a practical evaluation tool for financial LLMs and guides future development and optimization. The source code for Golden Touchstone and model weight of Touchstone-GPT have been made publicly available at https://github.com/IDEA-FinAI/Golden-Touchstone. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2411_06272 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Golden Touchstone: A Comprehensive Bilingual Benchmark for Evaluating Financial Large Language Models Wu, Xiaojun Liu, Junxi Su, Huanyi Lin, Zhouchi Qi, Yiyan Xu, Chengjin Su, Jiajun Zhong, Jiajie Wang, Fuwei Wang, Saizhuo Hua, Fengrui Li, Jia Guo, Jian Computation and Language Computational Engineering, Finance, and Science As large language models (LLMs) increasingly permeate the financial sector, there is a pressing need for a standardized method to comprehensively assess their performance. Existing financial benchmarks often suffer from limited language and task coverage, low-quality datasets, and inadequate adaptability for LLM evaluation. To address these limitations, we introduce Golden Touchstone, a comprehensive bilingual benchmark for financial LLMs, encompassing eight core financial NLP tasks in both Chinese and English. Developed from extensive open-source data collection and industry-specific demands, this benchmark thoroughly assesses models' language understanding and generation capabilities. Through comparative analysis of major models such as GPT-4o, Llama3, FinGPT, and FinMA, we reveal their strengths and limitations in processing complex financial information. Additionally, we open-source Touchstone-GPT, a financial LLM trained through continual pre-training and instruction tuning, which demonstrates strong performance on the bilingual benchmark but still has limitations in specific tasks. This research provides a practical evaluation tool for financial LLMs and guides future development and optimization. The source code for Golden Touchstone and model weight of Touchstone-GPT have been made publicly available at https://github.com/IDEA-FinAI/Golden-Touchstone. |
| title | Golden Touchstone: A Comprehensive Bilingual Benchmark for Evaluating Financial Large Language Models |
| topic | Computation and Language Computational Engineering, Finance, and Science |
| url | https://arxiv.org/abs/2411.06272 |