ChineseHarm-Bench: A Chinese Harmful Content Detection Benchmark
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866918124202229760 |
|---|---|
| author | Liu, Kangwei Cheng, Siyuan Tian, Bozhong Liang, Xiaozhuan Yin, Yuyang Han, Meng Zhang, Ningyu Hooi, Bryan Chen, Xi Deng, Shumin |
| author_facet | Liu, Kangwei Cheng, Siyuan Tian, Bozhong Liang, Xiaozhuan Yin, Yuyang Han, Meng Zhang, Ningyu Hooi, Bryan Chen, Xi Deng, Shumin |
| contents | Large language models (LLMs) have been increasingly applied to automated harmful content detection tasks, assisting moderators in identifying policy violations and improving the overall efficiency and accuracy of content review. However, existing resources for harmful content detection are predominantly focused on English, with Chinese datasets remaining scarce and often limited in scope. We present a comprehensive, professionally annotated benchmark for Chinese content harm detection, which covers six representative categories and is constructed entirely from real-world data. Our annotation process further yields a knowledge rule base that provides explicit expert knowledge to assist LLMs in Chinese harmful content detection. In addition, we propose a knowledge-augmented baseline that integrates both human-annotated knowledge rules and implicit knowledge from large language models, enabling smaller models to achieve performance comparable to state-of-the-art LLMs. Code and data are available at https://github.com/zjunlp/ChineseHarm-bench. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_10960 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | ChineseHarm-Bench: A Chinese Harmful Content Detection Benchmark Liu, Kangwei Cheng, Siyuan Tian, Bozhong Liang, Xiaozhuan Yin, Yuyang Han, Meng Zhang, Ningyu Hooi, Bryan Chen, Xi Deng, Shumin Computation and Language Artificial Intelligence Cryptography and Security Information Retrieval Machine Learning Large language models (LLMs) have been increasingly applied to automated harmful content detection tasks, assisting moderators in identifying policy violations and improving the overall efficiency and accuracy of content review. However, existing resources for harmful content detection are predominantly focused on English, with Chinese datasets remaining scarce and often limited in scope. We present a comprehensive, professionally annotated benchmark for Chinese content harm detection, which covers six representative categories and is constructed entirely from real-world data. Our annotation process further yields a knowledge rule base that provides explicit expert knowledge to assist LLMs in Chinese harmful content detection. In addition, we propose a knowledge-augmented baseline that integrates both human-annotated knowledge rules and implicit knowledge from large language models, enabling smaller models to achieve performance comparable to state-of-the-art LLMs. Code and data are available at https://github.com/zjunlp/ChineseHarm-bench. |
| title | ChineseHarm-Bench: A Chinese Harmful Content Detection Benchmark |
| topic | Computation and Language Artificial Intelligence Cryptography and Security Information Retrieval Machine Learning |
| url | https://arxiv.org/abs/2506.10960 |