Chinese SimpleQA: A Chinese Factuality Evaluation for Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Yancheng, Li, Shilong, Liu, Jiaheng, Tan, Yingshui, Wang, Weixun, Huang, Hui, Bu, Xingyuan, Guo, Hangyu, Hu, Chengwei, Zheng, Boren, Lin, Zhuoran, Liu, Xuepeng, Sun, Dekai, Lin, Shirong, Zheng, Zhicheng, Zhu, Xiaoyong, Su, Wenbo, Zheng, Bo
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912117571977216
author He, Yancheng
Li, Shilong
Liu, Jiaheng
Tan, Yingshui
Wang, Weixun
Huang, Hui
Bu, Xingyuan
Guo, Hangyu
Hu, Chengwei
Zheng, Boren
Lin, Zhuoran
Liu, Xuepeng
Sun, Dekai
Lin, Shirong
Zheng, Zhicheng
Zhu, Xiaoyong
Su, Wenbo
Zheng, Bo
author_facet He, Yancheng
Li, Shilong
Liu, Jiaheng
Tan, Yingshui
Wang, Weixun
Huang, Hui
Bu, Xingyuan
Guo, Hangyu
Hu, Chengwei
Zheng, Boren
Lin, Zhuoran
Liu, Xuepeng
Sun, Dekai
Lin, Shirong
Zheng, Zhicheng
Zhu, Xiaoyong
Su, Wenbo
Zheng, Bo
contents New LLM evaluation benchmarks are important to align with the rapid development of Large Language Models (LLMs). In this work, we present Chinese SimpleQA, the first comprehensive Chinese benchmark to evaluate the factuality ability of language models to answer short questions, and Chinese SimpleQA mainly has five properties (i.e., Chinese, Diverse, High-quality, Static, Easy-to-evaluate). Specifically, first, we focus on the Chinese language over 6 major topics with 99 diverse subtopics. Second, we conduct a comprehensive quality control process to achieve high-quality questions and answers, where the reference answers are static and cannot be changed over time. Third, following SimpleQA, the questions and answers are very short, and the grading process is easy-to-evaluate based on OpenAI API. Based on Chinese SimpleQA, we perform a comprehensive evaluation on the factuality abilities of existing LLMs. Finally, we hope that Chinese SimpleQA could guide the developers to better understand the Chinese factuality abilities of their models and facilitate the growth of foundation models.
format Preprint
id arxiv_https___arxiv_org_abs_2411_07140
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Chinese SimpleQA: A Chinese Factuality Evaluation for Large Language Models
He, Yancheng
Li, Shilong
Liu, Jiaheng
Tan, Yingshui
Wang, Weixun
Huang, Hui
Bu, Xingyuan
Guo, Hangyu
Hu, Chengwei
Zheng, Boren
Lin, Zhuoran
Liu, Xuepeng
Sun, Dekai
Lin, Shirong
Zheng, Zhicheng
Zhu, Xiaoyong
Su, Wenbo
Zheng, Bo
Computation and Language
New LLM evaluation benchmarks are important to align with the rapid development of Large Language Models (LLMs). In this work, we present Chinese SimpleQA, the first comprehensive Chinese benchmark to evaluate the factuality ability of language models to answer short questions, and Chinese SimpleQA mainly has five properties (i.e., Chinese, Diverse, High-quality, Static, Easy-to-evaluate). Specifically, first, we focus on the Chinese language over 6 major topics with 99 diverse subtopics. Second, we conduct a comprehensive quality control process to achieve high-quality questions and answers, where the reference answers are static and cannot be changed over time. Third, following SimpleQA, the questions and answers are very short, and the grading process is easy-to-evaluate based on OpenAI API. Based on Chinese SimpleQA, we perform a comprehensive evaluation on the factuality abilities of existing LLMs. Finally, we hope that Chinese SimpleQA could guide the developers to better understand the Chinese factuality abilities of their models and facilitate the growth of foundation models.
title Chinese SimpleQA: A Chinese Factuality Evaluation for Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2411.07140