INSEva: A Comprehensive Chinese Benchmark for Large Language Models in Insurance

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Shisong, Zhu, Qian, Yang, Wenyan, Yang, Chengyi, Wang, Zhong, Wang, Ping, Lin, Xuan, Xu, Bo, Li, Daqian, Yuan, Chao, Qi, Licai, Xu, Wanqing, zhenxing, sun, Lu, Xin, Xiong, Shiqiang, Chen, Chao, Hu, Haixiang, Xiao, Yanghua
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908519842709504
author Chen, Shisong
Zhu, Qian
Yang, Wenyan
Yang, Chengyi
Wang, Zhong
Wang, Ping
Lin, Xuan
Xu, Bo
Li, Daqian
Yuan, Chao
Qi, Licai
Xu, Wanqing
zhenxing, sun
Lu, Xin
Xiong, Shiqiang
Chen, Chao
Hu, Haixiang
Xiao, Yanghua
author_facet Chen, Shisong
Zhu, Qian
Yang, Wenyan
Yang, Chengyi
Wang, Zhong
Wang, Ping
Lin, Xuan
Xu, Bo
Li, Daqian
Yuan, Chao
Qi, Licai
Xu, Wanqing
zhenxing, sun
Lu, Xin
Xiong, Shiqiang
Chen, Chao
Hu, Haixiang
Xiao, Yanghua
contents Insurance, as a critical component of the global financial system, demands high standards of accuracy and reliability in AI applications. While existing benchmarks evaluate AI capabilities across various domains, they often fail to capture the unique characteristics and requirements of the insurance domain. To address this gap, we present INSEva, a comprehensive Chinese benchmark specifically designed for evaluating AI systems' knowledge and capabilities in insurance. INSEva features a multi-dimensional evaluation taxonomy covering business areas, task formats, difficulty levels, and cognitive-knowledge dimension, comprising 38,704 high-quality evaluation examples sourced from authoritative materials. Our benchmark implements tailored evaluation methods for assessing both faithfulness and completeness in open-ended responses. Through extensive evaluation of 8 state-of-the-art Large Language Models (LLMs), we identify significant performance variations across different dimensions. While general LLMs demonstrate basic insurance domain competency with average scores above 80, substantial gaps remain in handling complex, real-world insurance scenarios. The benchmark will be public soon.
format Preprint
id arxiv_https___arxiv_org_abs_2509_04455
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle INSEva: A Comprehensive Chinese Benchmark for Large Language Models in Insurance
Chen, Shisong
Zhu, Qian
Yang, Wenyan
Yang, Chengyi
Wang, Zhong
Wang, Ping
Lin, Xuan
Xu, Bo
Li, Daqian
Yuan, Chao
Qi, Licai
Xu, Wanqing
zhenxing, sun
Lu, Xin
Xiong, Shiqiang
Chen, Chao
Hu, Haixiang
Xiao, Yanghua
Computation and Language
Insurance, as a critical component of the global financial system, demands high standards of accuracy and reliability in AI applications. While existing benchmarks evaluate AI capabilities across various domains, they often fail to capture the unique characteristics and requirements of the insurance domain. To address this gap, we present INSEva, a comprehensive Chinese benchmark specifically designed for evaluating AI systems' knowledge and capabilities in insurance. INSEva features a multi-dimensional evaluation taxonomy covering business areas, task formats, difficulty levels, and cognitive-knowledge dimension, comprising 38,704 high-quality evaluation examples sourced from authoritative materials. Our benchmark implements tailored evaluation methods for assessing both faithfulness and completeness in open-ended responses. Through extensive evaluation of 8 state-of-the-art Large Language Models (LLMs), we identify significant performance variations across different dimensions. While general LLMs demonstrate basic insurance domain competency with average scores above 80, substantial gaps remain in handling complex, real-world insurance scenarios. The benchmark will be public soon.
title INSEva: A Comprehensive Chinese Benchmark for Large Language Models in Insurance
topic Computation and Language
url https://arxiv.org/abs/2509.04455