INSEva: A Comprehensive Chinese Benchmark for Large Language Models in Insurance
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908519842709504 |
|---|---|
| author | Chen, Shisong Zhu, Qian Yang, Wenyan Yang, Chengyi Wang, Zhong Wang, Ping Lin, Xuan Xu, Bo Li, Daqian Yuan, Chao Qi, Licai Xu, Wanqing zhenxing, sun Lu, Xin Xiong, Shiqiang Chen, Chao Hu, Haixiang Xiao, Yanghua |
| author_facet | Chen, Shisong Zhu, Qian Yang, Wenyan Yang, Chengyi Wang, Zhong Wang, Ping Lin, Xuan Xu, Bo Li, Daqian Yuan, Chao Qi, Licai Xu, Wanqing zhenxing, sun Lu, Xin Xiong, Shiqiang Chen, Chao Hu, Haixiang Xiao, Yanghua |
| contents | Insurance, as a critical component of the global financial system, demands high standards of accuracy and reliability in AI applications. While existing benchmarks evaluate AI capabilities across various domains, they often fail to capture the unique characteristics and requirements of the insurance domain. To address this gap, we present INSEva, a comprehensive Chinese benchmark specifically designed for evaluating AI systems' knowledge and capabilities in insurance. INSEva features a multi-dimensional evaluation taxonomy covering business areas, task formats, difficulty levels, and cognitive-knowledge dimension, comprising 38,704 high-quality evaluation examples sourced from authoritative materials. Our benchmark implements tailored evaluation methods for assessing both faithfulness and completeness in open-ended responses. Through extensive evaluation of 8 state-of-the-art Large Language Models (LLMs), we identify significant performance variations across different dimensions. While general LLMs demonstrate basic insurance domain competency with average scores above 80, substantial gaps remain in handling complex, real-world insurance scenarios. The benchmark will be public soon. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_04455 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | INSEva: A Comprehensive Chinese Benchmark for Large Language Models in Insurance Chen, Shisong Zhu, Qian Yang, Wenyan Yang, Chengyi Wang, Zhong Wang, Ping Lin, Xuan Xu, Bo Li, Daqian Yuan, Chao Qi, Licai Xu, Wanqing zhenxing, sun Lu, Xin Xiong, Shiqiang Chen, Chao Hu, Haixiang Xiao, Yanghua Computation and Language Insurance, as a critical component of the global financial system, demands high standards of accuracy and reliability in AI applications. While existing benchmarks evaluate AI capabilities across various domains, they often fail to capture the unique characteristics and requirements of the insurance domain. To address this gap, we present INSEva, a comprehensive Chinese benchmark specifically designed for evaluating AI systems' knowledge and capabilities in insurance. INSEva features a multi-dimensional evaluation taxonomy covering business areas, task formats, difficulty levels, and cognitive-knowledge dimension, comprising 38,704 high-quality evaluation examples sourced from authoritative materials. Our benchmark implements tailored evaluation methods for assessing both faithfulness and completeness in open-ended responses. Through extensive evaluation of 8 state-of-the-art Large Language Models (LLMs), we identify significant performance variations across different dimensions. While general LLMs demonstrate basic insurance domain competency with average scores above 80, substantial gaps remain in handling complex, real-world insurance scenarios. The benchmark will be public soon. |
| title | INSEva: A Comprehensive Chinese Benchmark for Large Language Models in Insurance |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2509.04455 |