Chinese SimpleQA: A Chinese Factuality Evaluation for Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912117571977216 |
|---|---|
| author | He, Yancheng Li, Shilong Liu, Jiaheng Tan, Yingshui Wang, Weixun Huang, Hui Bu, Xingyuan Guo, Hangyu Hu, Chengwei Zheng, Boren Lin, Zhuoran Liu, Xuepeng Sun, Dekai Lin, Shirong Zheng, Zhicheng Zhu, Xiaoyong Su, Wenbo Zheng, Bo |
| author_facet | He, Yancheng Li, Shilong Liu, Jiaheng Tan, Yingshui Wang, Weixun Huang, Hui Bu, Xingyuan Guo, Hangyu Hu, Chengwei Zheng, Boren Lin, Zhuoran Liu, Xuepeng Sun, Dekai Lin, Shirong Zheng, Zhicheng Zhu, Xiaoyong Su, Wenbo Zheng, Bo |
| contents | New LLM evaluation benchmarks are important to align with the rapid development of Large Language Models (LLMs). In this work, we present Chinese SimpleQA, the first comprehensive Chinese benchmark to evaluate the factuality ability of language models to answer short questions, and Chinese SimpleQA mainly has five properties (i.e., Chinese, Diverse, High-quality, Static, Easy-to-evaluate). Specifically, first, we focus on the Chinese language over 6 major topics with 99 diverse subtopics. Second, we conduct a comprehensive quality control process to achieve high-quality questions and answers, where the reference answers are static and cannot be changed over time. Third, following SimpleQA, the questions and answers are very short, and the grading process is easy-to-evaluate based on OpenAI API. Based on Chinese SimpleQA, we perform a comprehensive evaluation on the factuality abilities of existing LLMs. Finally, we hope that Chinese SimpleQA could guide the developers to better understand the Chinese factuality abilities of their models and facilitate the growth of foundation models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2411_07140 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Chinese SimpleQA: A Chinese Factuality Evaluation for Large Language Models He, Yancheng Li, Shilong Liu, Jiaheng Tan, Yingshui Wang, Weixun Huang, Hui Bu, Xingyuan Guo, Hangyu Hu, Chengwei Zheng, Boren Lin, Zhuoran Liu, Xuepeng Sun, Dekai Lin, Shirong Zheng, Zhicheng Zhu, Xiaoyong Su, Wenbo Zheng, Bo Computation and Language New LLM evaluation benchmarks are important to align with the rapid development of Large Language Models (LLMs). In this work, we present Chinese SimpleQA, the first comprehensive Chinese benchmark to evaluate the factuality ability of language models to answer short questions, and Chinese SimpleQA mainly has five properties (i.e., Chinese, Diverse, High-quality, Static, Easy-to-evaluate). Specifically, first, we focus on the Chinese language over 6 major topics with 99 diverse subtopics. Second, we conduct a comprehensive quality control process to achieve high-quality questions and answers, where the reference answers are static and cannot be changed over time. Third, following SimpleQA, the questions and answers are very short, and the grading process is easy-to-evaluate based on OpenAI API. Based on Chinese SimpleQA, we perform a comprehensive evaluation on the factuality abilities of existing LLMs. Finally, we hope that Chinese SimpleQA could guide the developers to better understand the Chinese factuality abilities of their models and facilitate the growth of foundation models. |
| title | Chinese SimpleQA: A Chinese Factuality Evaluation for Large Language Models |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2411.07140 |