InsQABench: Benchmarking Chinese Insurance Domain Question Answering with Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ding, Jing, Feng, Kai, Lin, Binbin, Cai, Jiarui, Wang, Qiushi, Xie, Yu, Zhang, Xiaojin, Wei, Zhongyu, Chen, Wei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929681935106048
author Ding, Jing
Feng, Kai
Lin, Binbin
Cai, Jiarui
Wang, Qiushi
Xie, Yu
Zhang, Xiaojin
Wei, Zhongyu
Chen, Wei
author_facet Ding, Jing
Feng, Kai
Lin, Binbin
Cai, Jiarui
Wang, Qiushi
Xie, Yu
Zhang, Xiaojin
Wei, Zhongyu
Chen, Wei
contents The application of large language models (LLMs) has achieved remarkable success in various fields, but their effectiveness in specialized domains like the Chinese insurance industry remains underexplored. The complexity of insurance knowledge, encompassing specialized terminology and diverse data types, poses significant challenges for both models and users. To address this, we introduce InsQABench, a benchmark dataset for the Chinese insurance sector, structured into three categories: Insurance Commonsense Knowledge, Insurance Structured Database, and Insurance Unstructured Documents, reflecting real-world insurance question-answering tasks.We also propose two methods, SQL-ReAct and RAG-ReAct, to tackle challenges in structured and unstructured data tasks. Evaluations show that while LLMs struggle with domain-specific terminology and nuanced clause texts, fine-tuning on InsQABench significantly improves performance. Our benchmark establishes a solid foundation for advancing LLM applications in the insurance domain, with data and code available at https://github.com/HaileyFamo/InsQABench.git.
format Preprint
id arxiv_https___arxiv_org_abs_2501_10943
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle InsQABench: Benchmarking Chinese Insurance Domain Question Answering with Large Language Models
Ding, Jing
Feng, Kai
Lin, Binbin
Cai, Jiarui
Wang, Qiushi
Xie, Yu
Zhang, Xiaojin
Wei, Zhongyu
Chen, Wei
Computation and Language
Artificial Intelligence
The application of large language models (LLMs) has achieved remarkable success in various fields, but their effectiveness in specialized domains like the Chinese insurance industry remains underexplored. The complexity of insurance knowledge, encompassing specialized terminology and diverse data types, poses significant challenges for both models and users. To address this, we introduce InsQABench, a benchmark dataset for the Chinese insurance sector, structured into three categories: Insurance Commonsense Knowledge, Insurance Structured Database, and Insurance Unstructured Documents, reflecting real-world insurance question-answering tasks.We also propose two methods, SQL-ReAct and RAG-ReAct, to tackle challenges in structured and unstructured data tasks. Evaluations show that while LLMs struggle with domain-specific terminology and nuanced clause texts, fine-tuning on InsQABench significantly improves performance. Our benchmark establishes a solid foundation for advancing LLM applications in the insurance domain, with data and code available at https://github.com/HaileyFamo/InsQABench.git.
title InsQABench: Benchmarking Chinese Insurance Domain Question Answering with Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2501.10943